Skip to main content

Kog Unlocks Deeper GPU Inference Potential

While Cerebras garnered significant market attention with its purpose-built AI chips following its May IPO, French startup Kog is pursuing a different

4 min read15 views5 tags
Originally reported bytechcrunch

While Cerebras garnered significant market attention with its purpose-built AI chips following its May IPO, French startup Kog is pursuing a different strategy, asserting that substantial untapped potential remains within conventional GPUs for accelerating AI inference.

Kog recently captured widespread interest on Hacker News in May with a technical preview. This demonstration aimed to validate its core premise: that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own,” showcasing its capabilities on widely used hardware like AMD MI300X and NVIDIA H200 GPUs.

Although some users expressed disappointment that this optimization did not extend to laptop GPUs, others quickly recognized its immense potential. As AI inference speed and cost emerge as critical bottlenecks, Kog's promise to unlock new capabilities on existing hardware through sophisticated software optimization drew considerable interest. "We had 200 tangible business leads," CEO Gaël Delalleau shared with TechCrunch.

Based on initial feedback, the solo founder anticipates that software engineering will be the primary early use case. Experienced users of tools like Claude Code are familiar with waiting hours for results, underscoring the value of speed. Anthropic itself acknowledges this, charging a price multiple for Claude’s Fast Mode.

Kog aims to attract customers who are currently hindered by such delays, particularly those relying on AI workflows for critical professional tasks. The startup is also collaborating with design partners who enable users to generate games and applications via prompts, for whom a faster outcome, facilitated by the Kog Inference Engine (KIE), directly translates into increased revenue, Delalleau noted.

The company acknowledges that the market is still evolving. Through its observations of demand, Kog discovered that prospective customers are not yet prepared to fine-tune smaller models. "And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen," Delalleau explained.

Achieving its ambitious promise of “30x faster LLM inference” represents a significant challenge for Kog. While its demonstration impressively showcased 3,000 per-request tokens per second (TPS), this was achieved using a custom-built small model with approximately 2 billion parameters, the now open-sourced Laneformer 2B.

Despite skepticism, Delalleau remains confident that their approach will be equally effective with large language models (LLMs), whose substantial size often poses a challenge for inference chips. "GPUs have a bright future," he stated. For Kog’s CEO, the notion that GPUs are ill-suited for decoding is a misconception, highlighting that newer GPUs offer increasing memory bandwidth that simply needs to be fully exploited.

Kog is not alone in recognizing the power of software optimization to enhance GPU performance beyond standard specifications. ZML, another French company, has developed hardware-agnostic software that bypasses Nvidia’s CUDA to enable fast inference across various competing chips. However, Delalleau positions Kog as more akin to Stanford University’s Hazy Research lab, emphasizing an even deeper-level focus on GPU acceleration.

Delalleau himself is not a researcher; his first startup, Stribe, a TechCrunch50 2009 alum, bears no relation to Kog, apart from his former co-founder, Kamel Zeroual, now a VC whose firm, Varsity VC, co-led Kog’s seed round. Nevertheless, the startup's profound technical focus originates from Delalleau's distinctive professional background.

After studying solid-state physics at France’s École Polytechnique, Delalleau transitioned into offensive cybersecurity, often referred to as white hat hacking. According to him, this diverse experience shaped the unique mindset he encourages within his team. From the scientific perspective, "there’s this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them."

Regarding his hacking experience, as a four-time finalist at DEF CON’s CTF tournament, Delalleau noted it taught him "to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn’t necessarily designed."

The intensity of this hands-on approach presents a challenge, as it is inherently time-consuming. "For every new GPU, we’ll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware." With a team of 11, this currently limits the number of chips Kog can optimize in the foreseeable future.

Longer-term, Kog plans to integrate its methodology into agent-based pipelines, which will enable support for a broader range of chips and models. This strategic direction aligns with Europe's ambition to develop its own capabilities in these areas, potentially providing sovereignty tailwinds for the startup, which already receives support from Scaleway and is backed by France’s Bpifrance and the French Tech 2030 program.

For the immediate future, however, Kog must demonstrate the efficacy of its approach on LLMs. This validation is also crucial for securing additional funding. "Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A," Delalleau concluded.

#AI News#Kog#GPU Inference#LLM Inference#Software Optimization
ES
Editorial StaffEditor

The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.

View all posts
Reader feedback

What did you think of this story?

User Comments

Filter:
No comments yet. Be the first to comment!
Continue reading
View all news