Our Mission
Open source software for data science benefits millions of practitioners all over the world, every day. We’re legally bound to keep it that way.
We’re a mission-driven, for-profit company
“When we founded Probabl, we chose to write our commitment to open source into our bylaws, not just our pitch deck. Scikit-learn became the foundation of machine learning because it was open, reliable, and built by a community of experts. Our job is to keep it that way for years to come, as data science moves from notebooks to production systems and AI agents.”

Yann Lechelle
Cofounder and Executive President of Probabl
Probabl is what is known as a mission-driven company in France (entreprise à mission) – a legal designation under Article L. 210-10 of the French Commercial Code that binds a company to pursue a social or environmental mission alongside its commercial activities.
The mission we inscribed in our bylaws is: “To develop, maintain at the state-of-the-art, and sustain a complete suite of open source tools for data science to benefit … the world.”
Specifically, we invest in the development, maintenance, and sustainability of the following open source tools:

Scikit-learn
The Python library for machine learning. Downloaded over 200 million times per month, over 5 billion times in total.
TabICL
The state-of-the-art fully open source tabular foundation model.
Skore
The Python library to evaluate and get insights from predictive models. It structures and stores your experiments so that you can easily retrieve them later.
Skrub
The Python library to ease preprocessing and feature engineering for tabular machine learning.

Skops
The Python library to share scikit-learn-based models and put them in production safely.
Every open source contributor we employ, every line of code they maintain, and every grant we pursue is in service of a single goal: keeping these open source tools for data science world-class and freely accessible to everyone.
Why this mission matters
Scikit-learn is the Python library for machine learning. It’s used to teach machine learning in classrooms and bootcamps across the world, and used in research labs and production systems across every major industry, from manufacturing and energy to healthcare and financial services. More recently, it’s used by agents, too.
In short, scikit-learn and the tools in its ecosystem are the open digital infrastructure of data science and machine learning that millions of practitioners depend on every day.
But software of this scale doesn’t maintain itself. Behind every release is a team of maintainers and contributors fixing bugs, reviewing pull requests, updating documentation, and keeping pace with a field that moves fast. Without sustained investment, even the most widely used tools in the world quietly decay.
Probabl exists because we believe that infrastructure this important deserves a steward – one with the resources, the expertise, and the legal obligation to keep it running.
5B+
Downloads in total
200M+
Downloads every month
1.3M+
Dependent GitHub repositories
28.7K+
Dependent open source packages
130K+
Academic citations
7,300+
Nature publications
How we’re held accountable to
our mission
Five experts. One mandate:
hold us to our mission.
A mission written into law requires oversight. Ours is provided by our mission alignment committee, which consists of one company representative and four external domain experts.
The committee reviews our progress and produces an annual Mission Alignment Report submitted to the French government. They have full access to our financial information and direct access to our Board. They can contest our findings, request additional evidence, and flag concerns publicly.
We’re proud to have the following experts on our inaugural Mission Alignment Committee.
Peter Wang
Chief AI and Innovation Officer and Co-Founder of Anaconda Inc.

“Machine learning has become central to every aspect of human life, and trusted, innovative open source projects like Scikit-Learn are a crucial part of that foundation.”
Emily Omier
Positioning consultant for open source companies and co-founder of Open Source Founders Summit

“It’s important for data science tools to remain as independent and transparent as possible; keeping them open source is the best way to do so.”
Mark Surman
President of Mozilla

“Open source creates choice. This is no different in data science and AI. With open source tools like scikit-learn and any-llm, data scientists can train, compare, and choose the best models for their use cases. Teams can swap models without rebuilding from scratch. Enterprises can switch vendors without starting over. That chain of freedom runs from the individual data scientist all the way to the top – but only if the foundations stay open and it takes dedicated people to develop, maintain, and sustain those foundations. That’s what Probabl exists to do.”
Arnaud Le Hors
Senior Technical Staff Member (Open Technologies – Open Source Security and AI) at IBM

“In an era where we see some companies pull the rug on open source communities by suddenly switching to a business license, it is great to see projects like scikit-learn that so many depend on in the data science and AI space being backed by companies like Probabl that are committed to staying on the open source path.”
Cailean Osborne, PhD
Head of Ecosystem Development at Probabl

“Open source is the bedrock of data science and AI. Python libraries like scikit-learn, pytorch, and many more are used daily by millions in labs, enterprises, and public institutions all over the world. But open source projects like these and above all the communities that underpin them don’t just sustain themselves. Their continued development requires dedicated stewardship and investment. Probabl was founded with this conviction, and that’s why I’m proud to be part of the team.”