French AI startup Kog is working to dramatically accelerate large language model (LLM) inference on conventional datacenter GPUs — without specialized hardware — and expects to demonstrate a first major model running at 10x speed as early as September 2026.
The Paris-based company, founded by CEO Gaël Delalleau, gained attention in May 2026 when a tech preview hit the front page of Hacker News, showing 3,000 tokens per second (TPS) on standard AMD MI300X and Nvidia H200 GPUs using a purpose-built 2-billion-parameter model called Laneformer 2B, which has since been open-sourced. The demo generated more than 200 business leads, according to Delalleau.
Kog’s core claim is that GPUs are underutilized for AI inference, and that deep software optimization — rather than purpose-built chips — can unlock significantly faster decoding speeds. The company’s Kog Inference Engine (KIE) targets a long-term goal of 30x faster LLM inference. Delalleau argues that newer GPUs carry increasing memory bandwidth that current software stacks fail to fully exploit.
The startup’s initial focus is on software engineering workflows, where AI tools like Claude Code can take hours to return results, and on platforms that let users generate games and apps from prompts. Faster inference in those contexts could translate directly to more revenue for design partners, Delalleau said.
Kog’s approach involves spending weeks or months reverse-engineering each new GPU at a low hardware level — a method Delalleau traces to his background in solid-state physics at École Polytechnique and offensive cybersecurity, including four finalist appearances at DEFCON’s CTF competition. With a team of 11, the depth of that process limits how many chips the company can support at once.
The seed round was co-led by Varsity VC, whose managing partner Kamel Zeroual previously co-founded a company with Delalleau. Additional backers include French public investment bank Bpifrance and cloud provider Scaleway. Delalleau said a Series A raise is planned once customer traction can be demonstrated following the September milestone.
Source: TechCrunch