Skip to benchmark
SPARKBENCHby LexiLominite
Loading public snapshot… Download data

Explore the run

Find your model.

Search the measured configurations, inspect every trial, or select up to three for a direct comparison.

More filters
Output speed uses each engine’s recorded timing basis.
CompareModelPrecisionEngineOutput speedFirst tokenResult

Side by side

Compare the trade-offs.

Output speed, response time, and whole-request speed answer different questions. Use matching engines for the fairest comparison.

Read the evidence

What the numbers mean.

A small, repeatable test of real model configurations on one DGX Spark. The caveats are part of the result.

A warmup, then three trials

One model runs at a time. A 16-token warmup comes before three measured requests capped at 256 generated tokens. Temperature is zero and the seed is fixed. SparkBench shows the median.

Three trials are a useful first check, not a confidence interval. Test order, temperature, clocks, and background activity can change a result.

Engine timing methods differ

Ollama output speed comes from the engine’s generated-token count and evaluation duration. vLLM output speed is estimated from the interval between the first and last streamed text.

Both are useful within their own configuration. Cross-engine differences include measurement method, tokenizer, kernels, and memory settings.

Context and caching matter

Ollama and the additional vLLM tests use an 8,192-token context. The existing FP8 service uses its 262,144-token configuration. Prompt caching was not disabled.

Prompt-processing speed can include cache reuse. Precision labels come from the recorded checkpoint or engine metadata; unknown stays unknown.

Generation is separate from embeddings

Embedding models process input rather than generate an answer, so SparkBench does not rank them beside generation models. Completion counts include reasoning tokens when an engine returns them.

GPT-OSS returned reasoning despite the request to disable it. Its result includes those tokens and does not represent visible-answer-only speed.

Show the public protocol and host snapshot

Full statistical report · Download aggregate summary

SparkBench · measured on NVIDIA DGX SparkPublic snapshot · inspectable methodology · no live machine controls