Pod Efficiency Analyzer

Watch a memory fabric that anticipates, and what it does to your margins.

returning session · “Given everything in this repository, where is the race condition?”
- time to first token token throughput
SESSION KV CACHE
~26× larger than GPU memory
INFERRA · MEMORY MOVEMENT SCHEDULER HBM ↔ DRAM ↔ NVMe
GPU · HBMscarce · compute-adjacent
GPU · ATTENTION KERNELSTALL
EPHEMERAL~80 GB · 1×
HOST DRAMwarm working set
STAGEDTB-class · ~25×
NVMe FABRICevery KV block · RDMA
PERSISTENTPB-scale · ~1000×
    328× faster

Calculator 01

Total Cost of Ownership

Your fleet, live cloud pricing, and what the same load costs with Inferra.

Capex spread over the hardware's life, plus power & colo opex, the effective $/node/hr for a neo-cloud.
Datacenter power
Cluster memory tiers
KV fabric total
Calculated from the workload you described and what this system can serve back.
Current compute / month
With Inferra / month (incl. license)
Monthly savings with Inferra
TCO reduction
Without Inferra
With Inferra

Calculator 02

Inference Profit Calculator

Your production cost per million tokens, and how low you can price while staying profitable.

Model details
vLLM's block_size. Cache is allocated whole blocks, and prefix sharing only matches on block boundaries.
vLLM's enable_prefix_caching. With it off nothing is reused between requests at all.
vLLM's own default is 0.90; the rest covers activations, workspace and fragmentation.
Token pricing
What you would list at. The market for this exact model is below.
Derived from your workload: demand ÷ fleet decode capacity.
Long-sequence tokens bill at a premium; only you can serve them.
Misses burn extra GPU; pass it through. Hits stay at list price.
Calibrate with your own measurements the fastest way to shrink the error band
Current model confidence ±36%
One real benchmark run collapses four separate error terms at once: achieved bandwidth, effective batch, prefill interference and framework overhead. This is by far the largest single improvement available.
From your own request logs: what share of prompt tokens are a repeated prefix. Drives the whole cache advantage, so measuring it is worth more than any other single number here.
Real traffic is not one context length. A spread of 3 means your p95 request is three times your median.
Recorded with your configuration so a proof of concept can run against the same build. It does not narrow the band on its own: we hold no per-version throughput data, so naming a version cannot honestly change the estimate.
If your stack speculates and we do not model it, we understate your decode throughput, potentially by more than half.
Why this matters more than more modelling. The four largest error terms are all things a single measurement resolves simultaneously. Adding further sub-models to an uncalibrated model tends to make it worse, not better, because every new term carries its own parameter error. Measurement collapses them.

Agent workloadwhat your customers run
How many of each, times what each one generates, gives the load on the fleet. The dashed column is the product, not something to fill in. This mix sets utilization, the calculated reuse share, and the animation's roster.
Fleet size, pricing, and license term carry over from the TCO Calculator tab. Reuse share and max context are calculated from the system you specified.
Tokens/month without Inferra
Tokens/month with Inferra
TCO per 1M tokens without
TCO per 1M tokens with Inferra
Margin per 1M output tokens (with Inferra)
Gross margin / month (with Inferra)
Lowest viable list price (30% margin)
Margin verdict at your price
Why this wins on OpenRouter: marketplaces rank providers on price, latency, uptime, and context length. Inferra cuts your cost per token so you can undercut, keeps turn-2 TTFT interactive so your latency stats shine, and lets you list context tiers competitors can't serve.

Evidence

Benchmark Results

Every number the calculators use, measured, sourced, reproducible. New results land here.

Turn-2 TTFT, returning session, warm cache LOG SCALE

Llama-4-Scout-17B-16E, 8×H100, FP8 KV, measured on Inferra, reproducible recipe

100KTOKENS
7.4 s
191 ms
39×
1MTOKENS
2 m 03 s
1.66 s
74×
10.5MTOKENS
1 h 42 m
18.7 s
328×

Also measured: DeepSeek-R1-70B @ 102K context 53.4 s → 275 ms (194×); 8×H200 10M-token warm TTFT 11.6 s (601×); GTC demo config (Mode-5 VP) 8.3 h → 26 s (1,154×) at 10M. All results bit-identical to full recompute; reuse survives engine and fabric restarts.

Extended context advantage

 

HBM alonethe GPU on its own
+ host DRAMstaged, microseconds away
+ NVMe fabricthe whole address space
drawn to scale 
more context addressable on the same node, because the cache is paged instead of pinned
The pitch to your customers: most serving stacks cap interactive multi-turn context around 32K–128K tokens. With Inferra you can market full-document and agentic sessions at 1M tokens and beyond, and on Llama-4-Scout, whose architecture supports a 10-million-token context window, KV-cache offload is what makes those sessions practical to serve: an entire large codebase, a year of logs, or a long-lived agent's accumulated memory, analyzed in one conversation, restored warm in seconds on any node. A listing no cache-regen competitor can match.

Published turn-2 TTFT measurements

4× NVIDIA L40S, ScaleFlux PCIe 5.0 NVMe, RDMA 800G - Lightbits Labs study · FarmGPU independent study

ModelContextTTFT without Inferra TTFT with InferraSpeedup

More results are landing here.

8×H200 ladders, vLLM KV-Connector runs, and per-model token-economics sweeps are queued for this page. Want your configuration measured? Ask for a PoC →

Deliverable

Generate Your PDF Report

One click compiles the boardroom PDF from everything you've dialed in.

First page of the generated report

The report compiles with live pricing and downloads right here. We keep your address so Lightbits Labs can follow up about a proof of concept. We also record how this site was used, and the organisation your network belongs to, to understand who the tool is helping. Nothing is sold or shared. Ask us at info@lightbitslabs.com to see or delete what we hold.

The server re-computes every figure with live Infracost pricing at generation time, so the PDF cites its own price source.

Prefer real numbers? Run the local analyzer

$ curl -fsSL https://api.pod-efficiency.tools/sbin/gpu-analyzer.py -o gpu-analyzer.py \ && python3 gpu-analyzer.py --pop --sign --session

Prefer to read it first? Download from api.pod-efficiency.tools/sbin/gpu-analyzer.py, then run python3 gpu-analyzer.py. Add --no-upload to keep results local only.

Reads hostname, CPU, RAM, GPUs and NFS mounts only · ~300 lines of auditable stdlib Python · uploads one JSON over HTTPS (--no-upload keeps everything local).

Next step

Talk to Lightbits Labs

Turn these estimates into a proof of concept, the button pre-fills an email with your numbers.

Sales & general inquiries

Email: info@lightbitslabs.com

USA & Canada: 1-866-614-9802

International: +1-408-547-4391

Or use the form on lightbitslabs.com/contact-us.

Email sales with my numbers

Opens your mail client with a summary of your fleet, workload profile, license term, and estimated savings, nothing is sent until you hit send.

Offices

USA: 1830 The Alameda, San Jose, CA 95126

Israel: 17 Atir Yeda St., Kfar Saba 4464313

What to ask for: an Inferra proof of concept on your workload, formal per-GPU license pricing for your commitment term, and sizing guidance for the NVMe cache layer behind your fleet.

How this works

How we work out these numbers

Plain English, first

This tool is a simulation. It works out what an inference cluster would cost and earn by calculating it from your specification, the same way you would on paper, just faster and with fewer arithmetic slips.

It is not a measurement of your cluster, and it is not a quote. Some parts of it are close to exact, like how much memory a model needs. Other parts are genuine estimates, like how fast your servers will actually run. We are open about which is which, and the tool shows you an honest range rather than pretending to a precision it does not have.

We work hard to keep this accurate and we keep improving it. Even so, our margin of error may be off, and real results will differ. By continuing to use the site you are looking at our best-effort simulation on that understanding. Please check anything important against your own testing before you spend money on it.

That is the whole message. Everything below is optional detail, organised so you can open only what you need.

Open by default, competitors included

This site is meant to be a public, free reference on inference economics, not a brochure. The arithmetic is readable, the assumptions are named, and the registers below are built to hold measured results for any system, ours and everyone else's, under one schema so they can actually be compared.

We publish results that do not flatter us. The B200 efficiency constant on this page was cut after a public MLPerf result showed our roofline was 22% optimistic, and that correction is documented in the Proofs pane rather than quietly applied. A comparison only means something if the losing case can appear in it.

If you have measured numbers, for Inferra or against it, send them. They go in under the same schema, with attribution, unedited.

Prove it before you spend anything

The honest way to check every claim on this site is to run it on your own workload, and you can do that without buying any hardware. Ask us for a proof of concept on machines you already have, or on ours. Nothing here needs to be taken on trust.

There is also a free local probe under the PDF Report tab. It runs on your existing cluster, reports what your hardware and workload actually do, and those measured numbers feed straight back into this calculator, roughly halving its error band.

Verify first, buy second. If our numbers do not survive contact with your workload, we would rather you found that out on a proof of concept than after a purchase order.

What is in the rest of this page

Nine sections, all optional. Here is what each one is for and when it is worth your time.

Accuracy, layer by layer

Different parts of this tool deserve different amounts of trust. The memory arithmetic is close to exact. The financial projections are not. This section separates them so you know which numbers to lean on.

The actual arithmetic

The real equations behind each number, with the constants written out and their sources named. Open this if you want to check our working, or if you are trying to reproduce a figure in your own spreadsheet.

What we deliberately left out

Every model leaves things out. These are ours. Read this before you rely on a number, because a few of these can move real outcomes by more than the differences this tool reports.

What we know, guess, and cannot foresee

A plain sorting of every input into what we compute directly, what we are estimating, and what nobody can predict. If you are building a business case, the middle bucket is where measuring pays for itself fastest.

Data sources and dependencies

Which figures are live, which are templates, and what the tool is built with. Useful if you want to judge how current the pricing is, or replace a template with your own quote.

Read the code yourself

Every projection algorithm, extracted from the page that just produced your numbers, so it cannot differ from what ran. The fastest way to verify us is to read it.

Other tools, and when to prefer them

We are not the only way to model this, and for some questions we are not the best one. This names four alternatives, how each works, and exactly where each beats this tool.

Conventions and disclosures

The housekeeping that changes how you read a number: how we count tokens and memory, who we are and what our commercial interest is, and what we have not yet validated.

Tell us where we are wrong

If a constant is off or a mechanism is missing, we want to hear it. This is how to reach us, and what kind of correction helps most.

The detail

Open whichever sections matter to you. Each one starts by saying what it contains.

Accuracy, layer by layerHow confident we are in each part, and why

Different parts of this tool deserve different amounts of trust. The memory arithmetic is close to exact. The financial projections are not. This section separates them so you know which numbers to lean on.

Confidence degrades as you move from arithmetic toward the future. The bars below are our own assessment, not a statistical confidence interval.

Memory arithmetic

KV bytes per token, weight footprint, what fits in HBM. Deterministic given the model architecture and quantization.

±2%

Concurrency and context ceiling

How many tenants and how much context fit at once. Follows from the memory arithmetic plus vLLM allocator behaviour.

±10%

Decode throughput

Bandwidth roofline with per-part efficiency. Real kernels, schedulers and software versions move this substantially, which is why the band is derived from a named error budget and narrows to about ±13% once you supply one real measurement.

±36% → ±13%

Prefill time and TTFT

Quadratic fit to a published measurement ladder, extrapolated beyond the measured range at long context.

±2x at extremes

Cost and margin

Arithmetic is exact, but it inherits every input price you gave us and every throughput estimate above.

±25 to 50%

Competitive position and captured demand

Live listings are real. What share of traffic a given price wins is a directional guess with no demand model behind it.

directional
The actual arithmeticEvery formula, every constant, nothing hidden

The real equations behind each number, with the constants written out and their sources named. Open this if you want to check our working, or if you are trying to reproduce a figure in your own spreadsheet.

Open any of these to see exactly what the site computes. The same arithmetic runs in the browser, in the PDF generator, and in the reference implementation you can read in the source viewer.

KV cache bytes per token The number everything else depends on

Attention keys and values are cached per token, per layer. Two families are modelled separately because their footprints differ by more than an order of magnitude.

# Grouped-query attention (Llama, Qwen, Mistral, GPT-OSS) kv_per_token = 2 x layers x kv_heads x head_dim x dtype_bytes # Multi-head latent attention (DeepSeek V3/R1/V4, Kimi) kv_per_token = 576 x layers x dtype_bytes # Hybrid stacks (Qwen3.5 Mamba, Jamba) cache only some layers kv_per_token = 2 x kv_layers x kv_heads x head_dim x dtype_bytes

Where the constants come from: layer counts, KV head counts and head dimensions are read from each model's published config.json. The MLA figure of 576 is the compressed latent dimension plus the decoupled rotary dimension, which is what actually lands in the cache. Known gap: a handful of very recent models in the selector carry estimated architecture metadata because no config has been published yet. Those are the least reliable entries in the list.

What fits in memory Weights, activation reserve, and the cache budget left over
usable = hbm_per_gpu x gpus x gpu_memory_utilization weights = total_params x quant_bytes reserve = activation and framework overhead kv_budget = usable - weights - reserve concurrent_tenants = kv_budget / (kv_per_token x context_length) max_context = kv_budget / (kv_per_token x tenants)

Assumptions: gpu_memory_utilization defaults to 0.90, matching the vLLM default, and is adjustable under expanded model details. Weights are charged at full parameter count because weight streaming is not yet modelled. Quantization options are gated to what the selected GPU family actually supports in hardware, so FP4 does not appear on architectures that cannot execute it. Not modelled: allocator fragmentation, the page granularity of the block allocator at small batch, and any memory held by co-resident services such as a draft model or a guardrail classifier.

Decode throughput A bandwidth roofline, deliberately not a kernel simulator

Decode is memory-bound. Each generated token requires reading the active weights and the whole KV cache for the batch, so tokens per second follows from HBM bandwidth.

bytes_per_step = active_weights + batch x kv_per_token x context tokens_per_sec = batch x hbm_bandwidth x efficiency / bytes_per_step # achieved fraction of peak, by part. Estimates, not measurements. H100 0.85 H200 0.84 A100 0.80 B200 0.78 L40S 0.72 L4 0.68

The efficiency factor: no real kernel reaches peak bandwidth, and the shortfall is not uniform across parts. HBM parts with a mature kernel stack sustain the most, GDDR6 parts such as the L40S and L4 sustain noticeably less on this streaming pattern, and a brand new architecture sustains less of its own very high peak than a well-tuned older one. These sit inside the 0.65 to 0.85 band that published serving measurements occupy. They are estimates, not measurements, and replacing them with figures from your own fleet is the cheapest accuracy improvement available to you. Mixture-of-experts models charge only the experts actually activated per token, not the total parameter count. Getting this wrong understated MoE throughput by up to 27x in an earlier revision of this site. Not modelled: speculative decoding, chunked prefill overlap, disaggregated prefill and decode, tensor-parallel communication cost, or the scheduler's real batching behaviour under mixed request lengths.

Prefill time and time to first token A quadratic fit, extrapolated at the top end
prefill_seconds = (A x n + B x n2) x scale A = 7.4529e-5 # linear term, per prompt token B = 4.841221e-11 # quadratic term, attention against itself scale = kv_per_token / 98304 # the fit was measured on one model, so it is # rescaled by how heavy this model's cache is

Fitted to the published Llama-4-Scout-17B-16E measurement ladder on 8x H100 with FP8 KV. The quadratic term is what makes long context expensive and is the entire reason cache reuse matters commercially. The honest limit: the fit is anchored on measurements up to roughly 10M tokens on one specific hardware and software configuration. Applying it to a different model, a different GPU generation, or a context beyond the measured range is extrapolation, and at the extremes we would not be surprised by a factor of two in either direction.

What Inferra changes Amdahl over the whole turn, not a headline multiplier

Published Inferra speedups are turn-2 warm-cache time to first token, which accelerates prefill only. Applying that figure to a whole workload would be dishonest, so it is discounted twice: once because only the reused share of requests benefits, and once because decode is untouched.

M = 1 / ( (1 - s) + s / k ) s = share of requests that hit a warm cache k = published speedup on the accelerated portion

Worked consequence: a 328x published figure at a 40% reuse share yields a whole-turn multiplier of about 1.66x, not 328x. That is the number the calculators use. One nuance the formula above hides: the implementation does not use that closed form directly. It builds the full turn, prefill plus decode, accelerates only the reused prefill, and takes the ratio. Because decode time sits in both numerator and denominator, the result is always lower than the closed form suggests. You can read the exact code under mult in the source viewer. Sources for k: the Lightbits LightInferra study on 4x L40S with ScaleFlux, the independent FarmGPU study, and the Inferra GA measurement ladder on 8x H100 and 8x H200. Those rigs are not your rig. Not modelled: RDMA fabric contention when many nodes prefetch at once, NVMe read amplification, and the cost of a cache miss that has to traverse the fabric before falling back to recompute.

The weakest assumption on this page. Inferra's concurrency advantage comes from needing only part of each sequence resident in HBM rather than all of it. We currently model that resident portion as a flat 32,768 tokens per sequence. That figure is a standing estimate, not a measurement, and it is not a fixed window: the real system evicts on capacity watermarks with residency scored by access history, so the resident share varies with pool size and access pattern. Deriving it from the actual watermarks and cache counters is the single highest-priority correction to this model. Until then, treat the Inferra concurrency column as the least settled number the analyzer produces.
Where the reuse share comes from Derived from your workload, and different in each mode

Cache reuse is not a slider you set, it is derived from the workload profile you describe: how conversational the traffic is, how much shared system prompt and tool schema sits in front of every request, and how long a session lives. Agentic traffic reuses heavily because each step replays the whole prior trajectory. One-shot completions barely reuse at all.

The derived share then differs by mode, which is the point of the comparison. Off gets no prefix caching. Others get HBM-resident prefix caching, so reuse is real but bounded by what fits in GPU memory and evicted the moment it does not. Inferra extends residency into NVMe and DRAM behind RDMA, so the same workload retains a materially larger share. Not modelled: the actual prefix-cache hit rate of your specific corpus, which is the single input most worth measuring before making a decision. The local probe under the PDF Report tab is the fastest way to replace this estimate with a measurement.

Cost, price, and margin Exact arithmetic over inputs you supply
node_hr = capex_amortized + power_kw x pue x kwh_price + licence cost_per_1M = node_hr x nodes / (tokens_per_sec x 3600 / 1e6) margin = (list_price - cost_per_1M) x monthly_tokens

Assumptions: straight-line amortization over the term you select, zero residual value at end of life, and a fixed PUE you can edit. Not modelled: financing cost and cost of capital, taxes and depreciation treatment, cloud egress, support and staffing, spares and RMA, contractual volume discounts, currency movement, and the cost of unsold capacity during ramp. A real business case includes all of these and this tool includes none of them.

Competitive position Live listings are real, the demand projection is not

Prices shown for comparable providers are fetched live and are genuine current listings for the selected model, with provider names withheld. Your position among them is a simple rank.

The part to distrust: any projection of how much traffic a given price captures. There is no demand model behind it. Real routing depends on latency, throughput, uptime history, context ceiling, rate limits, brand, contractual relationships, and where a given aggregator's defaults point. Price is one input among many. Treat share figures as an ordering, never as a volume forecast. Listings are cached for up to six hours, so a very recent price change may not be reflected.

What we deliberately left outThe gaps, named honestly, and how big they are

Every model leaves things out. These are ours. Read this before you rely on a number, because a few of these can move real outcomes by more than the differences this tool reports.

A simulation is defined as much by what it leaves out. None of the following are represented anywhere in the arithmetic, and each can move real outcomes by more than the differences this site reports.

Network fabric contentionScheduler and queueing effects Tail latency and SLO violationsNode failure and maintenance windows Noisy-neighbour interferenceSpeculative decoding Chunked and disaggregated prefillModel load and cold start Autoscaling lagTokenization and detokenization cost Guardrails, retrieval, rerankingServing stack version drift NVMe wear and write amplificationYour real prefix-cache hit rate Cost of capital and financingTaxes and depreciation treatment Egress, support, staffing, sparesDemand elasticity Utilization ramp and unsold capacityContractual discounting
Aggregate margin of error. Taking the layers together, treat capacity and memory conclusions as accurate to roughly ±10%, performance conclusions to roughly ±36% uncalibrated and ±13% once calibrated against your own measurements, and any financial projection as carrying an error band of ±25% to ±50% or worse, widening the further out you project. The relative comparison between Off, Others and Inferra is more robust than any absolute figure, because the same assumptions apply to all three columns and largely cancel.
What we know, guess, and cannot foreseeSorted into three honest buckets

A plain sorting of every input into what we compute directly, what we are estimating, and what nobody can predict. If you are building a business case, the middle bucket is where measuring pays for itself fastest.

Before anyone commits capital on the back of a simulation, they should be able to say which category each number falls into.

Known knowns

Things we compute or observe directly. Errors here would be bugs, not uncertainty.

  • KV bytes per token for every model with a published config.
  • Weight footprint at each supported quantization.
  • Whether a configuration fits in the memory available.
  • Live competitor listings for the selected model.
  • Your own inputs: node count, power price, term, list price.
  • The published benchmark ladder and the rigs it ran on.

Known unknowns

Things we know we are estimating. These are where measurement pays for itself fastest.

  • Your real cache reuse share. The single highest-leverage unknown on this page.
  • Achieved bandwidth efficiency on your stack, against our per-part estimates of 0.68 to 0.85.
  • Prefill behaviour for your model on your hardware beyond the fitted range.
  • Speedup transfer from published rigs to your fabric and NVMe.
  • Architecture metadata for the newest models, some of it estimated.
  • Real utilization once traffic is uneven across the day.

Possible unknown unknowns

Things that have historically surprised people in this exact market. We cannot size these, which is the point.

  • A step change in model architecture that alters cache economics outright, as MLA already did once.
  • Hardware price movement. HBM, DRAM, NVMe and GPU pricing have all moved sharply and in both directions within single quarters.
  • Energy price and grid access shifting faster than contracts can absorb.
  • A competitor pricing below cost for strategic reasons, making rational pricing uncompetitive.
  • Serving stack releases that change throughput by large factors without warning.
  • Demand moving to different model sizes entirely, stranding a fleet specced for today's mix.
  • Regulatory or licensing change affecting where and how models may be served.
Data sources and dependenciesWhere the numbers and the software come from

Which figures are live, which are templates, and what the tool is built with. Useful if you want to judge how current the pricing is, or replace a template with your own quote.

Data sources, dependencies, and the parts of the system you can inspect for yourself.

Model metadata

Layer counts, attention geometry and parameter counts from published config.json files on Hugging Face, entered into the selector by hand.

Benchmark ladder

Lightbits LightInferra study (4x L40S with ScaleFlux), the independent FarmGPU study, and the Inferra GA measurement ladder on 8x H100 and 8x H200.

Live token pricing

Current per-model listings across hosted providers, refreshed on a six hour cache. Provider identities are deliberately withheld.

Hardware and cloud pricing

Infracost Cloud Pricing API for instance rates, plus template hardware pricing you can override with your own quotes.

Serving assumptions

vLLM defaults for gpu_memory_utilization, block_size, prefix caching and continuous batching behaviour.

The stack

Python and Flask on the server, hand-written JavaScript and SVG in the browser with no charting library, and pdflatex for the board report. No analytics or advertising third parties.

Read the code yourselfThe modelling algorithms, as they actually run

Every projection algorithm, extracted from the page that just produced your numbers, so it cannot differ from what ran. The fastest way to verify us is to read it.

Every projection and modelling algorithm described above is readable in your browser, annotated, with the assumptions marked inline. We would rather you check our arithmetic than take our word for it.

Open the source viewer
Other tools, and when to prefer themFour alternatives, and what each does better than us

We are not the only way to model this, and for some questions we are not the best one. This names four alternatives, how each works, and exactly where each beats this tool.

If you are sizing inference infrastructure seriously, do not rely on any single model, including this one. These take deliberately different approaches, and where two of them disagree with us, they are more likely to be right about systems behaviour than we are. We link them because triangulation is how you find out.

NVIDIA Dynamo DynoSim

Discrete event

Approach: event-driven simulation of the full request lifecycle. Its mocker core reproduces batching, prefix-cache hits, preemption and block allocation across a disaggregated prefill and decode topology, with separate timing models predicting compute duration. It also simulates router and planner scaling decisions.

Versus this tool: DynoSim models the behaviour we deliberately do not, above all the scheduler and the queue. If your question is how a serving topology behaves under real request arrival patterns, use DynoSim. If your question is what it costs, use this. It is the natural next step after this page.

docs.nvidia.com/dynamo · DynoSim

Vidur

Profiling based

Approach: profile the operators once on real GPUs, fit execution-time predictors from those measurements, then simulate any number of configurations without needing a GPU again. Covers tensor and pipeline parallelism, several schedulers, and synthetic or trace-driven workloads, reporting TTFT, time per output token and end-to-end latency. From Microsoft Research.

Versus this tool: Vidur is grounded in measurement where we are grounded in a roofline, which makes its performance numbers considerably more trustworthy than ours. It does not model money. Running Vidur to obtain a throughput figure and feeding that number into this calculator is a better workflow than trusting our estimate.

github.com/microsoft/vidur

LLMServingSim

HW/SW co-simulation

Approach: hardware and software co-simulation for heterogeneous and disaggregated serving. Accelerator compiler and simulator stacks plug in underneath, so non-GPU and mixed-silicon fleets can be explored. The authors report tracking real GPU serving systems to within 14.7% while running far faster than cycle-level accelerator simulators. From the CASYS group at KAIST.

Versus this tool: it reaches down to the accelerator and the interconnect, which is exactly the layer our per-part bandwidth efficiency estimates paper over. It is also the right tool if you are evaluating hardware that is not an NVIDIA GPU, which this site does not model at all.

github.com/casys-kaist/LLMServingSim

llm-analysis

Closed form analytical

Approach: closed-form arithmetic over model architecture, hardware specification, dtype and parallelism strategy to estimate latency and memory for training and inference, with no benchmarking required. It explicitly describes its output as a lower-bound estimate.

Versus this tool: methodologically our closest relative, and a good independent check on our memory and latency arithmetic specifically. It is more thorough than we are on parallelism strategies and recomputation. We add fleet economics and cache reuse, which it does not attempt.

github.com/cli99/llm-analysis

MLPerf Inference

Measurement, not simulation

Approach: a standardized, audited benchmark suite with defined scenarios and accuracy targets, run on real hardware by the vendors themselves and published for comparison. Not a simulator at all, which is precisely why it belongs on this list.

Versus this tool: every simulator on this page, ours included, is ultimately trying to predict something that MLPerf measures. If a published MLPerf result exists for a configuration close to yours, it beats all of our estimates and you should use it instead.

mlcommons.org · Inference (Datacenter)

Know of a simulator, benchmark harness or capacity planner that belongs beside these? Send it to us and we will add it, including tools that compete directly with ours.

Conventions and disclosuresUnits, token accounting, who built this, what is unvalidated

The housekeeping that changes how you read a number: how we count tokens and memory, who we are and what our commercial interest is, and what we have not yet validated.

Who built this

This tool is built and published by Lightbits Labs, the vendor of Inferra, which is one of the three options it compares. We have an obvious commercial interest in the Inferra column looking good.

That is exactly why the arithmetic is on this page and the source is readable. The comparison uses identical assumptions across all three columns so that shared errors cancel, and published speedups are discounted rather than quoted at headline value. Judge the method, not the vendor.

Token accounting

Prices and costs are expressed per 1M tokens using a single blended rate. Most commercial providers price input and output tokens separately, often at a ratio of three or four to one, so a blended comparison can flatter or penalise you depending on your real input to output mix.

If your workload is prefill-heavy, which agentic traffic usually is, treat our competitive comparison as approximate and re-check it against split pricing before setting a real list price.

Units

Memory is computed in binary units (GiB, 230 bytes). A card marketed as 80GB is treated as 80 GiB, which is deliberate: HBM stacks come in powers of two, so this tracks what nvidia-smi actually reports far better than reading the marketing figure as decimal would.

ECC and driver-reserved memory are not subtracted separately, they sit inside the utilization factor. Bandwidth is decimal GB/s as published. Currency is USD throughout, with no conversion or tax handling.

Validation status

Honest answer: we have not yet published a formal backtest. The prefill fit is anchored on the published ladder it was derived from, which is not independent evidence, and the roofline has not been systematically compared against held-out measurements.

That work is the top priority for this model, and results will be published here with residuals, including the ones that make us look bad. Until then, weight the confidence bars above accordingly.

Data freshness

Token listings refresh on a six hour cache. Cloud instance rates come from a live pricing API. Hardware capex, networking and storage prices are template figures carrying a review date, not live quotes, and hardware pricing has moved sharply in recent years.

Override them with your own quotes before drawing any conclusion about capital cost.

What we record

We record page interactions, the configurations specced here, coarse location derived from IP, and an email address if you choose to give one for the PDF. It is first-party only, with no advertising or third-party analytics networks.

We use it to understand which configurations people model and to follow up on sales enquiries. Ask us to delete your record at any time.

Tell us where we are wrongCorrections that change the numbers get applied

If a constant is off or a mechanism is missing, we want to hear it. This is how to reach us, and what kind of correction helps most.

If a constant looks off, an architecture entry is stale, or a whole mechanism is missing from the model, we want to hear it. Corrections that change the numbers get applied and credited.

Especially useful: measured throughput from your own fleet, a real prefix-cache hit rate, a hardware quote that contradicts our template pricing, or a published benchmark that contradicts our fit.

Suggest an improvement
Proofs: public benchmarks behind our baseline What we compare against, and how each number reaches our formulae

Our baseline is not a claim, it is arithmetic anchored on published third-party measurements. This lists each anchor, what it measured, and exactly which constant in our model it sets. Where our formula disagrees with a public result, the disagreement is stated with its sign and size rather than smoothed away.

MLPerf Inference v5.1 · Llama-2-70B

validates the decode roofline

Measured: 8×B200 reached 101,611 tok/s server and 101,246 tok/s offline. 8×H200 reached roughly 33,000 tok/s on the same model. Audited, vendor-submitted, reproducible.

How it reaches our model: we reproduce that configuration through our own decode roofline. Llama-2-70B at FP8 is 80 layers × 8 KV heads × 128 head dim, so 160 KB of cache per token; at batch 1024 and 2048 context the step reads about 385 GB. Against 8 × 8 TB/s that predicts 123,594 tok/s.

We were 22% optimistic. Server and offline agreeing to within 0.4% means the gap is not a latency constraint, it is achieved bandwidth. So this result now sets our B200 decode efficiency directly: 0.64, replacing an estimated 0.78. One public benchmark, one constant, no judgement left in it.

mlcommons.org · Inference (Datacenter)

Lightbits LightInferra study

sets k, the prefill speedup

Measured: 4×L40S with ScaleFlux NVMe. Turn-2 warm-cache TTFT 152× to 286× faster; throughput 19× at 100K context.

How it reaches our model: supplies k in the capacity multiplier. It is never applied at face value: k accelerates prefill only, and only on the reused share, so the whole-turn figure is far smaller than the headline.

Vendor-run, on a rig unlike a modern B200 fleet. The single least independent number in the model.

FarmGPU independent study

corroborates k

Measured: L40S with ScaleFlux and 800G RDMA. Up to 1,154× TTFT improvement at 10M context, 3× more requests on the same GPUs, 65% lower infrastructure cost.

How it reaches our model: a third-party check on the same mechanism. Its "3× more requests" is the closest public analogue to our capacity multiplier, which currently lands near 3.7× on the default configuration.

blog.farmgpu.com

Inferra GA measurement ladder

sets the prefill fit

Measured: Llama-4-Scout-17B-16E on 8×H100 with FP8 KV, turn-2 warm: 100K in 191 ms, 1M in 1.66 s, 10.5M in 18.7 s. Continuous batching with one user. Turn 1 is a cold time to first token; turn 2 is warm with virtual paging and streaming. At 10.5M the baseline could not run at all on that hardware: the cache did not fit, which is why the warm path needs paging rather than merely benefiting from it.

How it reaches our model: the cold side fits A and B, which set time to first token and whether a request meets its SLA. It is a one-user latency, so it is deliberately not used to size fleet capacity: charging a 64-way batch as though every request were served alone overstated the cost of prefill several times over. Capacity uses a FLOP bound instead, at a stated 40% MFU.

Still the weakest constant we have. One model, one rig, one software version, one user, scaled to other models by cache size. The MFU behind the capacity bound is an estimate too. A single measured prefill rate from a .pop would replace both, and would be the most valuable number anyone could send us.

What is missing, plainly. Three of the four anchors above concern the accelerated path, and only one is independent of us. We have no audited public measurement of Inferra on a current B200 fleet, and no published backtest of this model against held-out results. Until those exist, treat the Inferra column as the least validated part of the page.
Benchmark register How measured runs will validate the numbers on this site

A place to put real measured runs, from us and from anyone else, so that the assertions on this site can eventually be backed by a specific benchmark rather than by a formula. Not yet switched on.

The intent is to accept results from the common public harnesses, guidellm, vLLM's benchmark suite, GenAI-Perf and MLPerf-style runs, and store them beside the hardware and software they ran on. Every Inferra result carries the exact build it came from, so a claim on this page can be traced to a commit rather than to a marketing round number, and results from different builds are never silently averaged together.

# every registered run branch inferra branch name required commit full commit hash required committed_at commit date and time, UTC required build build number optional release release number optional harness guidellm | vllm-bench | genai-perf | mlperf hardware GPU model, count, interconnect software serving stack and version config model, quantization, batching, context result tok/s, TTFT p50/p95, achieved batch, cache hit rate

Results for competing systems are first-class entries, not footnotes: same schema, same required fields, same visibility, with the submitter recorded. The register is intended to be readable by anyone without an account, so a buyer can compare our hardware and our uplift against public results on other systems without taking our word for any of it, including results that do not flatter us.

Currently off. ValidateNumbersViaBenchmarksDB is false, so nothing on this site is validated against the register yet and no figure here depends on it. When it is enabled, assertions that a registered benchmark covers will show the run behind them, and assertions with no matching run will be marked as modelled rather than measured. Turning it on will not change any number, only what each number can prove about itself.
The .pop format Signed, verifiable performance attestations anyone can produce

A Performance Optimization Package is a signed container holding a benchmark run, the system it ran on, and the raw harness output it came from. It is how a result gets listed here without anyone having to trust the party who submitted it, including us.

The problem it solves is simple. A number in a slide deck is an assertion. A number in a .pop carries the hardware it ran on, the exact software build, every setting that could change the result, the unmodified harness output, and a signature over all of it. Change one byte and the digest stops matching. That makes a competitor's submission worth exactly as much as ours, which is the point.

What is inside

result.pop # zip container, deterministic ordering ├── manifest.json format version, producer, created_at, digest algo ├── system.json GPU model/count, interconnect, host, driver,firmware, NUMA and PCIe topology ├── config.json model, quantization, KV dtype, TP/PP degree,batching mode, context, feature flags, env ├── provenance.json branch, commit, committed_at, build, release ├── results/ RAW harness output, byte-for-byte unmodified │ ├── guidellm.json │ └── vllm-bench.json ├── metrics.json normalised to the site schema, derived from results/ └── attestation/ ├── digest.txt SHA-256 over every file above, path-sorted ├── signature.sig detached signature over digest.txt └── chain.pem signer certificate chain

Raw output is mandatory. metrics.json is only ever derived from results/, never hand-entered, so anyone can recompute the normalised numbers from the harness output and check we did not massage them.

Who can produce one

Inferra emits a .pop directly. Anything else needs only a compliant converter, which is a thin wrapper that collects system state and wraps existing harness output:

  • guidellm, vLLM benchmark suite, GenAI-Perf and MLPerf-style runs are the intended first-class inputs.
  • A converter must not alter results/. It may only add the surrounding metadata and compute the digest.
  • The format carries no vendor-specific field. A package describing a system with no Inferra in it is a valid package.

What happens on upload

  • Digest recomputed from the container contents and compared to digest.txt. Any mismatch is rejected outright.
  • Signature verified against the chain. A package that verifies is listed as attested; one that is well-formed but unsigned may still be listed, clearly marked unverified, because an unsigned real result beats no result.
  • Normalisation re-derived from results/ rather than trusting metrics.json, and a disagreement is shown rather than silently resolved.
  • Listed publicly under the three registers, readable without an account.
Not yet accepting uploads. The three registers are off, so no .pop is being ingested and no figure on this site depends on one. The format is documented now so that converters can be written against a stable target, and so that the schema is public before we start collecting anything under it. When it opens, it opens for everyone's hardware on the same terms.
Published packages What other people are running, measured and modelled, kept apart

Every package anyone chose to publish. Attested runs and simulated projections are listed separately and never ranked against each other: a modelled number alongside a measured one would make the measurement look like just another opinion. Load any row into the simulator, or take it as a .pop.

Loading…

Publication is opt-in, per package. Nothing you configure here is published unless you press publish, and nothing is collected silently. A submission carries the configuration and its headline numbers, never an identity.
Correctness register Does the accelerated path return the same tokens as the slow one

Speed is worthless if the answer changes. This defines where we record correctness runs: the exact configuration a result came from, and whether output stayed bit-identical to an uncached baseline under it. Not yet switched on.

A cache that is fast and subtly wrong is worse than no cache, and the failure is quiet: a changed logit shows up as a slightly different answer, not as an error. The register pins every knob that could change numerics, so a correctness claim always names the mode it was established in rather than being asserted for the product as a whole.

# configuration under test branch / commit / committed_at as the benchmark register build / release optional feature_flags VP+Streaming, prefetch on/off, ... prefetch_algo two-signal | temporal | similarity | none env environment variables affecting numerics settings watermarks, eviction policy, block size model_config model, quantization, KV dtype, TP degree # what was checked mode off | others | inferra token_match bit-identical | divergence rate vs baseline logit_delta max and mean absolute deviation eval task scores against the uncached run seeded whether sampling was pinned
Currently off. ValidateCorrectnessDB is false. No correctness claim on this site is currently backed by a stored run, and none of the performance numbers depend on it. Enabling it lets a stated uplift carry the correctness evidence for the exact flag combination that produced it.
Stability register Does it stay correct and fast for weeks, not minutes

Benchmarks run for minutes; fleets run for quarters. This defines where long-duration runs are recorded, including the failures, so stability is evidenced rather than assumed. Not yet switched on.

Nothing on this site currently models node failure, maintenance, memory growth or degradation over time, and a cache tier is exactly the kind of component whose behaviour drifts as it fills. Recording error and crash logs against the build that produced them is the only honest way to close that gap.

# run identity branch / commit / committed_at as the benchmark register build / release optional duration wall clock, and tokens served over it hardware / software as deployed # what happened throughput_drift tok/s at start vs end latency_drift TTFT p50 and p95 over time memory host and device growth, leak signature errors counts by class, with log references crashes failure logs, restarts, time to recover availability measured uptime and drain events
Currently off. ValidateStabilityDB is false. Note what this means for the figures above: node failure, maintenance windows and degradation are listed among the things this model does not represent, and they stay unrepresented until this register holds real runs.
Source code and credits Everything readable, and who built it

Where to read the code behind every number, the tools that feed it, and who is responsible for it.

Modelling algorithms

Every projection algorithm, extracted from the served page at the moment you load it, so it provably cannot differ from the code that produced your numbers.

open.pod-efficiency.tools/source

Local probe

The stdlib-only script you run on your own cluster to report what your hardware and workload actually do. Readable before you run it, as anything you pipe into a shell should be.

/sbin/gpu-analyzer.py

Full repository

The site, the server, the PDF generator and the infrastructure definitions live in one repository. It is currently private, so the link below will not open for you yet.

github.com/arthurrasmusson-lb/pod-efficiency-tools

Licence

Everything published here is released under Creative Commons Attribution 4.0 International (CC BY 4.0). Use it, change it, build on it, commercially or not. The only condition is attribution.

creativecommons.org/licenses/by/4.0

What is not published

The analytics and admin code, the infrastructure configuration, and commercial licence pricing are deliberately excluded from the in-browser viewer. Everything that shapes a performance, capacity or cost result is included.

If you want to check a number, the source viewer is the fastest route: it is generated from the running page rather than from a repository, so it cannot fall out of date with what you are looking at.

Licensed CC BY 4.0. Attribution is the only condition, so a fork, a rewrite, or a competing calculator built on this arithmetic is all fine by us, which is rather the point of publishing the working.

Created by

Claude Code Opus 5, Fable 5, and Arthur Rasmusson

Published by Lightbits Labs, the vendor of Inferra, which is one of the three options this tool compares. That commercial interest is disclosed in full under Conventions and disclosures.

Legal noticeWhat we are and are not responsible for

The formal version of the plain-English note at the top of this page. Worth reading before you use these figures in a business case.

Important notice regarding use of these projections

The Pod Efficiency Analyzer is a simulation and modelling tool provided for informational and illustrative purposes only. It does not constitute financial, investment, procurement, engineering or legal advice, and it is not an offer, quotation, or commitment of any kind. No performance level, cost, saving, margin, or competitive position shown here is warranted, guaranteed, or contractually promised.

All outputs are estimates generated from assumptions described on this page and from inputs supplied by you. They depend on published third-party benchmarks measured on specific hardware and software configurations that are unlikely to match your own environment, and on components of a real deployment that this model does not represent at all, including those listed above under what this model does not include.

Projections tied to financial outcomes are additionally subject to market demand for individual models, competitive pricing changes by other providers, and movement in the price of GPUs, HBM, DRAM, NVMe, networking, and electricity. These have historically changed materially and rapidly. Present prices shown or assumed here may not hold, and future changes may affect outcomes in either direction. Live pricing data is retrieved from third-party sources, may be cached, delayed, incomplete, or inaccurate, and is presented without verification.

To the maximum extent permitted by law, Lightbits Labs and the authors of this tool accept no liability for any loss or damage whatsoever, including without limitation lost profits, lost revenue, wasted expenditure, or procurement, capital allocation or pricing decisions taken in reliance on these outputs, arising from any source, including without limitation any inaccuracy in the modelling, any component not modelled, any change in prices, costs, demand, competition, technology or regulation, or any other cause.

Do not use this tool as the sole basis for a financial or investment decision. Validate against a proof of concept on your own workload and hardware, and against independent quotations, before committing capital.

Reopen the pre-simulation notice

Serve 3× more inference on the GPUs you already have.

Generate my PDF report Contact sales
Talk to sales →