.png)
.png)
.png)
Synthetic Data for AI agents
We simulate expert-level reasoning and domain-specific processes to generate training data that builds true specialization into your models. Synth handles the cold start problem, covers the long tail of edge cases through engineered simulations, and runs entirely on-premise to keep sensitive data under your control.
Learn moreFully Open Data for AI
We curate and structure the world's largest rights-cleared and provenance-based dataset for LLMs - government records, legal archives, scientific literature, multilingual sources - so you can plug it directly into your models, RAG pipelines, and MCP servers.
Learn moreAI-Native Tooling for Agents
Turn your messy siloed documents into a single, structured, compliant data asset for your agentic AI workflows. Your AI systems and agents get not only the right information but rich trustworthy context - higher accuracy on a wider range of processes. A built-in privacy firewall handles PII before anything leaves the secure zone. Deployable fully on-premise.
Learn more
BlogOn the occasion of the AI for Good Global Summit in Geneva, we are open-sourcing the models, datasets, and deployment frameworks presented in this post: the reasoning models framework for RAG, our on-device benchmarking harness for cache-augmented generation, and the French-language retrieval pipeline for Raspberry Pi and Android. Everything described here runs fully offline, on hardware costing less than €100.


BlogToday we introduce the Telco Common Corpus - 10B+ tokens of fully open telecommunications data, highlighting, as part of the GSMA Open Telco AI initiative, how the company is helping AI work harder for the telecoms sector.
BlogPleias trained a 600-million-parameter specialized model for RATP to detect and interpret safety signals in Parisian Subway users’ messages - combining a fully synthetic training pipeline and designed for on-premise deployment. After only three months of development, the model, beating closed models 200x times its size, is now in production at RATP’s sovereign infrastructure on Scaleway.
Use CasePleias trained a 600-million-parameter specialized model for RATP to detect and interpret safety signals in Parisian Subway users’ messages - combining a fully synthetic training pipeline and designed for on-premise deployment. After only three months of development, the model, beating closed models 200x times its size, is now in production at RATP’s sovereign infrastructure on Scaleway.
Use CaseResearchers built an AI medical assistant that works without internet for health workers in rural West Africa. It runs on old Android phones and handles local languages to give workers quick treatment guidance.
Use CasePleias and SpineDAO are partnering to build AI systems that safely scale expert spine care for back pain, the world's leading cause of disability. The project tests whether small, structured-reasoning models can outperform large generic LLMs in high-stakes clinical settings.



