
DeFi3 min read
Spark Vulnerability Mitigated
A vulnerability was identified via responsible disclosure and has been mitigated in the Spark protocol. It affects multi-input Spark spends only; single-input spends are unaffected. Your wall
BitcoinWorld Kog says software can unlock 30x faster LLM inference on existing GPUs French startup Kog is betting that software optimization can dramatically accelerate large language model i
BitcoinWorld
Kog says software can unlock 30x faster LLM inference on existing GPUs
French startup Kog is betting that software optimization can dramatically accelerate large language model inference on standard data center GPUs, claiming up to 30x speed gains without requiring new hardware. The company, which emerged from stealth in May 2025, says it has already attracted over 200 business leads, signaling strong enterprise demand for faster AI response times.
As AI models become integral to professional workflows, the time it takes for a model to generate a response has become a critical bottleneck. Slow inference not only hampers productivity but also increases operational costs, especially for applications like AI-assisted coding, where users may wait minutes or even hours for results. Anthropic, for example, charges a premium for its ‘Fast Mode’ on Claude, underscoring the value of speed.
Kog’s approach targets this pain point directly. The company’s technology, dubbed the Kog Inference Engine (KIE), is designed to optimize the decoding process on existing GPUs from NVIDIA and AMD, such as the H200 and MI300X. In a technical preview that reached the front page of Hacker News, Kog demonstrated 3,000 tokens per second on a single request using a small 2-billion-parameter model, which they open-sourced as Laneformer 2B.
While the initial demo used a small model, Kog’s CEO Gaël Delalleau says the company is now focused on scaling its techniques to large language models, a significant technical leap. “We’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen,” Delalleau told Bitcoin World. The company has learned that most potential customers are not interested in fine-tuning small models for their specific tasks, making support for mainstream LLMs essential.
Delalleau is confident that the same optimization principles will work with larger models, despite skepticism from some quarters. He argues that modern GPUs have abundant memory bandwidth that is underutilized by current inference software. “GPUs have a bright future,” he said, rejecting the notion that they are ill-suited for decoding.
Kog is not alone in exploring software-based inference acceleration. ZML, another French startup, has developed hardware-agnostic software that bypasses NVIDIA’s CUDA to support fast inference across competing chips. However, Delalleau draws a distinction, comparing Kog’s work to Stanford’s Hazy Research lab, with a deeper focus on GPU-level engineering.
This deep-level focus stems from Delalleau’s background in solid-state physics and offensive cybersecurity. He explains that this combination of understanding the laws of hardware and reverse-engineering at the assembly level allows his team to push GPUs beyond their intended performance envelopes. But this approach is time-consuming; for each new GPU, the team dedicates weeks or months to detailed engineering research.
Kog’s initial market focus is on software engineering, where slow inference is a known pain point. The company also has design partners in the app and game generation space, where faster output translates directly to revenue. However, the market for ultra-fast inference is still maturing, and Kog is adapting its roadmap accordingly.
The company, which has a team of 11, is backed by French institutions including Bpifrance and the French Tech 2030 program, and is supported by cloud provider Scaleway. Delalleau anticipates that demonstrating a major model running at 10x speed, possibly by September 2025, will be key to securing a Series A round. “Once we’ve implemented our first major model at 10x speed, we’ll be able to start demonstrating customer traction,” he said.
Kog’s promise to unlock more inference performance from existing GPUs addresses a pressing need in the AI industry. If successful, it could offer a cost-effective alternative to specialized hardware, giving enterprises a way to accelerate AI workloads without significant capital expenditure. The coming months will be critical for the startup to prove its technology at scale.
Q1: What is Kog’s core technology?Kog’s Kog Inference Engine (KIE) is a software layer that optimizes the decoding process of large language models on standard data center GPUs, aiming to deliver up to 30x faster inference speeds compared to conventional methods.
Q2: How does Kog achieve faster inference?Kog uses deep-level GPU engineering techniques, including low-level reverse engineering and hardware-aware optimizations, to better utilize the memory bandwidth and compute capabilities of existing GPUs from NVIDIA and AMD.
Q3: When will Kog’s technology be available for large models?Kog is currently focused on scaling its approach to large language models. The company expects to demonstrate a major model running at 10x speed by September 2025, with broader availability likely following a successful Series A round.
This post Kog says software can unlock 30x faster LLM inference on existing GPUs first appeared on BitcoinWorld.