AI startup Probably raised $9 million in seed funding in 2026 to develop a system designed to eliminate hallucinations and factual errors from AI-generated outputs, with backing from Andreessen Horowitz.
Founded by Peter Elias, Probably is building what he describes as a “data science mech suit” — an engineering framework that wraps a large language model (LLM) in a deterministic validator system. When the LLM produces an answer, the validator checks it against the underlying dataset and rejects any result that doesn’t match. The LLM has also been trained against the validator, with the full system optimized for speed and accuracy.
The company’s first product is a data science tool that generates answers from complex datasets, with each result accompanied by a citation and an audit trail. The stated goal is to reach 99.99% accuracy — the kind of reliability common in deterministic systems but rarely achieved with AI.
One practical consequence of the approach is that it can run on significantly smaller models. Elias says the current version uses a model “four classes weaker than the frontier models,” which means it can run on local hardware rather than data center infrastructure, reducing token costs.
“What we learned building this was that the better your harness engineering is, the weaker the model can be,” Elias said. “If you can refine the context enough, the model does not have to work very hard to do the right thing.”
Elias says the same underlying engine could extend beyond data science to fields such as accounting and medical services — what he calls “any precision-sensitive use case.”
The approach may carry broader relevance at a time when, as the company notes, token costs are rising and customers are reassessing AI spending. Elias also suggested that major AI labs have little incentive to solve the hallucination problem themselves: “They make money the more times you have to correct the model.”
Source: TechCrunch