Paris-based AI startup ZML released ZML/LLMD in July 2026, a free LLM inference server designed to run open-source large language models across a wide range of chips — including those from Nvidia, AMD, Google (TPU), Apple (Metal), and Intel Arc.
The software aims to eliminate vendor lock-in by allowing enterprises and cloud providers to mix and match chips at maximum available speed, and in some cases faster, according to ZML founder Steeve Morin. “The idea is to give people back the power to create their own system and achieve real efficiency gains,” Morin told TechCrunch.
ZML/LLMD is not open source, unlike the company’s earlier ML inference framework released in 2024 and updated in March. It is launching as a free product while the company studies usage patterns before deciding if and where to introduce paid tiers. Morin said he would rather measure adoption before monetizing than risk constraining growth too early.
The release comes as AI inference — the processing of user prompts — has grown in strategic importance relative to model training, attracting what Morin described as an “inference gold rush.” Competitors in the space include Baseten, recently valued at $13 billion, Inferact from the creators of vLLM, and RadixArk, the commercial entity behind SGLang.
Morin said ZML’s scope extends beyond inference software into co-designing silicon with chipmakers, several of which are European, including Axelera, Fractile, Kalray, and SiPearl, among others.
ZML operates with a team of 20 and has raised $20 million from investors including 20VC, Kima Ventures, LocalGlobe, and Kindred Capital. Notable angel investors include Docker and Dagger founder Solomon Hykes, Hugging Face co-founders Clément Delangue and Julien Chaumond, and Turing Award winner Yann LeCun, now with AMI Labs.
Morin, a former VP of engineering at Zenly — acquired by Snapchat for nine figures in 2017 — said the startup’s Paris base is central to what it is building. “I couldn’t do ZML anywhere but in Paris,” he said.
Source: TechCrunch