The takeaway
Software-level optimization, not new silicon, may be the fastest route to cheaper, lower-latency AI inference.
Why it matters for builders
For builders running agentic workloads, Kog's bet signals that GPU inference is far from optimized. Software-level acceleration — monokernels, delayed tensor parallelism, and continuous weight streaming — could unlock 10x-plus speedups on hardware teams already own, cutting token costs and latency for real-time agents without a hardware upgrade.
French Startup Kog Squeezes More AI Inference Out of GPUs
The race for faster AI inference has a new contrarian entrant. Markets rewarded Cerebras and its purpose-built chips with a warm IPO debut in May, but French startup Kog is betting that conventional GPUs still have far more performance locked inside them than anyone has managed to extract.
Software over silicon
Kog hit the front page of Hacker News in May with a tech preview aimed at proving that "extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own" — namely the AMD MI300X and Nvidia H200. Its Kog Inference Engine (KIE) hit roughly 3,000 output tokens per second on a single request, though the demo ran on a purpose-built 2-billion-parameter model, Laneformer 2B, which the startup has since open-sourced.

That demo translated into real demand. "We had 200 tangible business leads," CEO Gaël Delalleau told TechCrunch. The first target use case is software engineering, where veteran Claude Code users "sometimes have to wait hours to get results."
The 30x promise
Kog's headline ambition is "30x faster LLM inference," but the gap between a 2B-parameter model and frontier-scale LLMs is substantial. The company has refocused entirely on accelerating larger models, since prospective customers aren't willing to fine-tune small ones.
Delalleau rejects the pessimism around GPUs for decoding. "GPUs have a bright future," he said, pointing to ever-increasing memory bandwidth that "only begs to be unlocked." The tradeoff is a labor-intensive approach: Kog's 11-person team spends "several weeks or even months" reverse-engineering each new GPU, a mindset Delalleau traces to his background in solid-state physics and DEFCON-level offensive security.
Why it matters
Kog isn't alone. Fellow French startup ZML released hardware-agnostic software that bypasses Nvidia's CUDA, while Stanford's Hazy Research lab pushes in a similar direction. Kog's deeper, GPU-specific focus is what sets it apart, and it arrives as inference speed becomes the industry's new battleground.
Backed by Bpifrance, French Tech 2030, and Scaleway, Kog is also riding European AI-sovereignty tailwinds. The company plans to demonstrate its first major model running at 10x speed by September — a milestone that would unlock its Series A. For builders watching inference costs climb, that's a signal worth tracking.
The Automation Brief
Read 5 AI stories instead of 50.
The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.
No noise. Unsubscribe anytime.
Editorial notes
Stefan Trbojevic
n8n Lab Editorial
15 August 2026
15 August 2026
AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.


