Wafer raised a $40 million Series A on September 1 to automate more of the engineering required to run AI models quickly and cheaply. Marathon Management Partners and Chemistry co-led the round, with Wing, AMD Ventures and Outset Capital participating alongside returning investors Fifty Years and Y Combinator, according to Wafer’s announcement.

The funding follows a $4 million seed disclosed four and a half months earlier and brings Wafer’s announced financing to $44 million. RuntimeWire reported that the company initially planned to raise $18 million. The much larger round gives Wafer room to pursue a wider technical and commercial target: continuous optimization across the inference stack.

For a buyer, Wafer’s claim should be tested against that buyer’s workloads and service constraints. A provider can post an excellent throughput result under a carefully chosen model, accelerator, batch size and request profile. Production systems face their own arrival rates, prompt lengths, output lengths, concurrency spikes, latency requirements and availability targets. Wafer’s product thesis makes those workload details part of the optimizer’s input, so customers should make them part of the purchasing test as well.

From generated kernels to a running control loop

Wafer began in 2025 with an agent that writes and optimizes GPU kernels, the low-level programs that govern how effectively a workload uses accelerator hardware. It has since moved further up the serving stack. RuntimeWire, citing Wafer’s Y Combinator profile, says Wafer’s agents modify kernels, batching, scheduling and memory layouts for open-source models. Wafer’s own announcement separately names the model, inference engine, kernels and hardware as optimization targets.

The distinction is architectural. Kernel tuning improves individual pieces of execution. Full-stack tuning can search across interacting choices: which model variant to deploy, which serving engine should run it, which kernels fit the hardware, and how the deployment should respond to observed traffic and performance constraints.

A locally faster kernel can still sit inside a poorly matched deployment. Larger batches may raise aggregate throughput while extending the wait experienced by an individual request. A memory layout may work well for one model and accelerator pairing while constraining another. Hardware with a weaker software ecosystem may become economically attractive after enough work on kernels and serving configuration. Wafer’s proposed loop treats these choices as a connected search problem.

Production traffic can shift after deployment.

A configuration selected during a predeployment benchmark reflects the workload measured at that moment. Wafer says its system learns from traffic patterns and performance constraints, then keeps finding ways to improve performance per dollar. If the system works as described, optimization becomes an operating process attached to the deployment rather than a finite engagement performed before launch.

The Series A will finance further automation of that loop. Wafer has shared the destination, though its funding announcement provides few details about the machinery: how often deployments are retuned, which changes happen automatically, how candidate configurations are tested, what rollback controls exist, or how the optimizer balances cost against latency and reliability. Those details will determine whether continuous optimization behaves like dependable infrastructure or an ambitious managed service.

What static leaderboards can and cannot show

Tokens per second remains useful. It measures something concrete, and it can eliminate configurations that plainly lack enough capacity. The number becomes weak evidence when treated as a purchasing verdict.

Wafer’s own pitch supplies the reason. If performance depends on workload traffic, model choice, engine, kernels and hardware, a benchmark that fixes those variables measures one point in a large configuration space. An operator’s production distribution may occupy another point entirely.

The procurement question should therefore become empirical: how does each provider run this model under our request distribution while meeting our latency, reliability and cost limits?

For an interactive assistant, the evaluation might weight time to first token, output cadence and tail latency under bursts. A background extraction service may tolerate slower individual requests in exchange for lower cost and higher aggregate throughput. An agent system introduces sequences of calls whose delays compound across a task. The correct weighting belongs to the application owner because the user experience and economic constraints belong to that application.

Static results still help build a shortlist. The contract decision deserves a replay of representative traffic, including ordinary load, peaks, long contexts, varied output lengths and whichever failure conditions the operator can safely induce. Providers should receive the same workload and the same service-level objectives. Cost should be measured at the level the buyer pays, then examined alongside latency distributions, error rates, availability and performance during load transitions.

That makes a single throughput score insufficient for a contract decision.

A better inference trial

Wafer’s continuous-optimization claim suggests a two-stage evaluation.

The first stage measures the starting deployment. Give each provider a fixed model or an explicitly bounded set of acceptable variants, replay the same request sample, and record the service indicators tied to the application. Include total cost for the trial, since Wafer sells performance per dollar as part of its proposition.

The second stage measures adaptation. Change the traffic distribution within a predefined envelope and allow each provider to tune its deployment. The operator should record how quickly performance recovers, which components change, whether the changes disturb live requests, and whether the resulting configuration still satisfies every service objective.

That second stage is the harder test for Wafer. Many providers can prepare a fast benchmark configuration. Wafer is raising $40 million around the claim that its software can keep improving a deployment as conditions evolve. A trial that freezes the workload tests the resulting configuration while barely touching the company’s central claim.

Operators also need a stable comparison boundary. Allowing one provider to swap models while another must preserve exact model behavior produces a muddy result. Model substitutions can alter quality, safety behavior or output consistency even when the API remains compatible. The trial should state which layers each provider may change and which application properties must remain fixed.

The same discipline applies to price. A lower cost per generated token can coexist with higher end-to-end expense if retries, longer outputs, capacity reservations or operational labor rise. A useful trial should weigh availability, reliability, latency under representative traffic, data-handling requirements and price stability alongside raw model speed. Put those requirements into the test before providers tune against them. The Signal’s earlier analysis of Fireworks AI’s inference platform examines the same buyer problem at the serving layer.

Continuous tuning creates a control problem

Automation introduces a governance question: who decides when a faster or cheaper configuration is safe to promote?

Wafer’s public announcement says the system learns from workload traffic and performance constraints. It leaves the promotion process unexplained. An operator should ask whether optimized configurations pass a shadow deployment, canary traffic or another controlled validation step before receiving production requests. The answer affects both reliability and accountability.

Constraint design becomes part of the product. A team that supplies a throughput target while omitting tail latency, quality checks or failure budgets may get an efficient deployment that violates an unstated requirement. Continuous optimizers can only search within the objective and boundaries they receive.

Change records matter too. When the optimizer modifies an engine, kernel, batch policy, memory layout or accelerator selection, the operator needs enough provenance to connect that change with subsequent behavior. Otherwise, a performance regression becomes an archaeological dig across a system specifically designed to keep changing.

A sensible deployment policy would separate safe tuning parameters from changes that require human approval. Batch sizing within a tested range may deserve one treatment. Switching accelerators, engines or model variants may deserve another, especially where deterministic behavior, regulated data or customer commitments constrain the serving environment. Wafer has yet to describe its policy surface publicly in either supplied source.

AMD’s strategic interest

AMD Ventures’ participation fits Wafer’s hardware-level ambition. Wafer runs workloads on Nvidia and AMD accelerators and plans support for additional architectures, according to RuntimeWire. The publication argues that years of software work around Nvidia’s CUDA stack have left competing accelerators with fewer mature optimizations.

Software capable of generating kernels and retuning serving configurations could help close that gap for specific workloads. The economic possibility is straightforward: an accelerator with an attractive rental price becomes more useful when software can recover performance that generic configurations leave unused.

Wafer said in July that it ran Z.ai’s GLM-5.2 on AMD MI355X accelerators at more than twice the performance per dollar of an Nvidia Blackwell deployment. RuntimeWire correctly labels the result as Wafer’s own testing. It supports further investigation, while an operator-grade conclusion requires the model configuration, workload distribution, latency targets, pricing assumptions and reproducible test procedure.

AMD Ventures also has a direct commercial interest in software that makes AMD hardware easier to deploy. Its participation supplies strategic alignment and capital. Independent workload trials must supply the buying evidence.

The proof Wafer still owes

Wafer’s fundraising announcement is concise by design. It states the round, investors, product direction and intended use of capital. RuntimeWire adds company history, the expansion from kernels into managed inference, the hardware strategy and several business claims.

Among those claims, RuntimeWire reports that Wafer reached roughly $8 million in annualized revenue within 12 weeks, based on comments from co-founder Emilio Andere. That is a short-period run rate, equivalent to about $666,000 at the reported monthly pace, rather than revenue collected across a full year. The publication also reports gross margins around 50 percent and identifies Vercel and Inworld AI as customers.

Those numbers indicate early demand for the service. They leave open the durability of revenue, the cost of workload-specific engineering and the amount of optimization performed autonomously. A system marketed as autonomous performance engineering earns stronger economics only when software absorbs work that would otherwise scale with engineers, customer deployments or rented compute.

The operational evidence should come in longitudinal form: performance and cost before optimization, the production change that triggered retuning, the configuration selected, time to improvement, effect on reliability, and results sustained across weeks or months. Comparable disclosures across Nvidia and AMD hardware would also test Wafer’s claim that optimization can move among accelerator choices instead of remaining tied to one favored stack.

Watch after September 1, 2026, for Wafer’s first reproducible production case showing how its optimizer reacts to a changed traffic distribution while preserving declared latency, reliability and quality constraints.


The Signal is the public edge of a private practice. Sherpa points the same intelligence engine at one owner's business — competitors, suppliers, regulators, watched daily, graded and sourced. Work with a Sherpa →