Analysis
A French startup with just 11 people is making one of the boldest efficiency claims in AI infrastructure this year: Kog says its software can unlock up to 30x faster large-language-model inference on the same GPUs enterprises already have installed, according to TechCrunch. No new hardware, no swapped-out chips -- just a different approach to how decoding runs on standard datacenter GPUs like Nvidia's H200 and AMD's MI300X.
The company, backed by French public investment bank Bpifrance and the government's French Tech 2030 program, and supported by cloud provider Scaleway, first drew attention in May when a technical preview hit the front page of Hacker News. The pitch, dubbed the Kog Inference Engine, is that 'extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own' -- a direct challenge to the industry's default assumption that faster inference requires newer, more expensive silicon.
Kog sits in a crowded and increasingly well-funded category tracked on Pulse's AI chip startup dashboard: GPU financiers have been pivoting toward inference-chip investment since July, and fellow French startup ZML has already released a free product aimed at speeding inference across multiple AI chip types. What differentiates Kog's approach, per its own claims, is depth: rather than optimizing at the scheduling or batching layer the way most inference-acceleration tools do, Kog says it goes deeper into the decoding process itself, which is where its 30x figure comes from on best-case workloads.
“That number needs real scrutiny before anyone treats it as representative.”
That number needs real scrutiny before anyone treats it as representative. Inference-speed multipliers this large are typically measured on narrow, favorable benchmarks -- single-request decoding on a specific model architecture -- and rarely hold up unchanged across the mixed, high-concurrency workloads that make up most production AI traffic. Kog's own roadmap acknowledges this gap: the company says its first implementation on a major model at a more modest 10x speed is expected in September, a meaningfully lower number than the headline 30x figure, and the company plans to use that milestone to demonstrate customer traction ahead of a Series A raise.
The timing lines up with a broader shift in where AI infrastructure capital is flowing. Anthropic's reported $6 billion talks to acquire Decart are built on a similar logic -- buying software that makes existing compute go further, rather than buying more chips outright -- and Tencent's own executives have said openly that renting out underused AI hardware can be more profitable than running models on it. Kog is a much smaller, earlier-stage bet on the same underlying thesis: that efficiency software, not raw GPU count, is where a meaningful chunk of AI's next cost savings will come from.
Kog's Series A will be the real test of whether investors believe the 30x claim translates into a fundable business -- expect the September model-implementation milestone to be the number that actually determines the round's size and valuation, not the headline figure from May's Hacker News preview.