Connection
On s2 this is the internal URL ending in /v1. Leave auth empty.
Settings sent to Ollama
CPU options for this host
16 vCPUs and a large RAM pool. These go in Ollama options. Change one value, run the speed check twice, and keep the second tokens/sec.
num_thread 16 uses every vCPU. If 16 is slower, try 12, then 8. num_batch 512 is the prompt batch; try 256 and 1024. num_gpu 0 leaves the model on CPU. use_mlock pins the weights in RAM. Leave numa off on a VM. keep_alive 30m keeps the model loaded between checks.
Sampling and cache
Leave a box empty to omit it. Ollama then uses its own default. These fields are the same keys a Modelfile PARAMETER line would set, and this page sends them on every REST call.
f16_kv on uses more RAM and is the usual cache. Off can be smaller and slower. seed makes two runs comparable. stop cuts the answer at a marker, which shortens generation.
Variants
One Ollama server, many request shapes. Save the form under a name, then run the speed-check prompt on each name. The calls go out one after another so load time shows when a second model tag has to enter RAM.
| name | model | gen tok/s | load |
|---|
Modelfile for the current form
Create this on the Ollama host when you want a permanent name. After that, REST calls use "model": "your-name". Options sent by this page still override the file for that one call.
Speed
Generation tok/s is answer tokens divided by Ollama’s generation time. It does not include loading the model. The first run loads weights into RAM. Compare the second run.
The check sends one short prompt and ignores the chat history, so each row is the same test.
| gen tok/s | prompt tok/s | load | threads | batch | ctx | model |
|---|
Throughput
Sends the speed-check prompt 1, 2, or 4 times at once. Total tok/s is every output token divided by wall-clock time. Per-call tok/s is the average of each call’s own generation rate.
If total tok/s stays flat when you go from 1 to 2, Ollama is queuing. Set OLLAMA_NUM_PARALLEL on the Ollama container to at least the number of calls in flight, then redeploy Ollama and run again. Parallel slots also multiply the KV cache, so keep num_ctx at 4096 while you test.
| in flight | total tok/s | per-call tok/s | output tokens | wall | model |
|---|
Pipeline content
Translation uses low reasoning and a 51,200-token window, enough for 32,000 characters of HTML and the same length of XML. Style extraction uses medium reasoning and a 131,072-token window. A 200,000-token document is split into chunks. The prompts are the translator files, filled for Portuguese to English (UK).
Chat
Last request and response
Send a message or click Probe.