Ollama chat tester

Change the settings, send a short message, and see the exact request that went to Ollama.

Connection

On s2 this is the internal URL ending in /v1. Leave auth empty.

Settings sent to Ollama

CPU options for this host

16 vCPUs and a large RAM pool. These go in Ollama options. Change one value, run the speed check twice, and keep the second tokens/sec.

num_thread 16 uses every vCPU. If 16 is slower, try 12, then 8. num_batch 512 is the prompt batch; try 256 and 1024. num_gpu 0 leaves the model on CPU. use_mlock pins the weights in RAM. Leave numa off on a VM. keep_alive 30m keeps the model loaded between checks.

Sampling and cache

Leave a box empty to omit it. Ollama then uses its own default. These fields are the same keys a Modelfile PARAMETER line would set, and this page sends them on every REST call.

f16_kv on uses more RAM and is the usual cache. Off can be smaller and slower. seed makes two runs comparable. stop cuts the answer at a marker, which shortens generation.

Variants

One Ollama server, many request shapes. Save the form under a name, then run the speed-check prompt on each name. The calls go out one after another so load time shows when a second model tag has to enter RAM.

name model gen tok/s load

Modelfile for the current form

Create this on the Ollama host when you want a permanent name. After that, REST calls use "model": "your-name". Options sent by this page still override the file for that one call.


    

Speed

Generation tok/s is answer tokens divided by Ollama’s generation time. It does not include loading the model. The first run loads weights into RAM. Compare the second run.

— generation tok/s
— prompt tok/s
— model load
— output tokens

The check sends one short prompt and ignores the chat history, so each row is the same test.

gen tok/s prompt tok/s load threads batch ctx model

Throughput

Sends the speed-check prompt 1, 2, or 4 times at once. Total tok/s is every output token divided by wall-clock time. Per-call tok/s is the average of each call’s own generation rate.

If total tok/s stays flat when you go from 1 to 2, Ollama is queuing. Set OLLAMA_NUM_PARALLEL on the Ollama container to at least the number of calls in flight, then redeploy Ollama and run again. Parallel slots also multiply the KV cache, so keep num_ctx at 4096 while you test.

in flight total tok/s per-call tok/s output tokens wall model

Pipeline content

Translation uses low reasoning and a 51,200-token window, enough for 32,000 characters of HTML and the same length of XML. Style extraction uses medium reasoning and a 131,072-token window. A 200,000-token document is split into chunks. The prompts are the translator files, filled for Portuguese to English (UK).

Chat

Last request and response

Send a message or click Probe.