PERFORMANCE, DOWN TO THE KERNEL

Every token.
Less waiting.

We build optimized kernels and faster inference systems. Bringing models closer to the limits of the hardware they run on.

Build something faster
THE DECODE RELAYCONCEPTUAL PIPELINE / 001
01 / INPUTModel + workload
02 / COMPUTEOptimized kernel
03 / OUTPUTNext token. Faster.
01 / OUR FOCUS

Small improvements.
At massive scale.

Inference is a chain of decisions. We work where better code and a deeper understanding of the hardware make every step count.

01

Kernel engineering

Purpose-built compute. Exploring fusion, memory access, and execution patterns to make the critical path more efficient.

02

Faster inference

From prefill to decode. Reducing the overhead between the model, the runtime, and the hardware to keep tokens moving.

03

Performance that holds up

Real workloads. Clear baselines. Measuring latency, throughput, and correctness together so improvements stand up beyond a microbenchmark.

02 / THE APPROACH

The hardware has limits.
Let’s get closer.

01

Find the bottleneck.

Start with the workload. Profile the path from input to output and understand where time goes.

02

Make the critical path count.

Optimize the work that matters: memory movement, compute, and the handoffs in between.

03

Measure. Then iterate.

Check correctness. Compare against a baseline. Keep the changes that deliver.

03 / START A CONVERSATION

Have a workload
worth accelerating?

hello@decoderelay.com

Let’s talk kernels, inference, and what faster could look like.