Estudos
Jul 28, 2026
Ollama Truncation Loop — or: How 4096 Context Tokens Cost Me 2 Hours of Debugging
Hermes kept entering a continuation loop 4 times before giving up. The cause: Ollama, by default, only allocates 4096 context tokens — regardless of the model supporting 262K. The hard way: num_ctx of the model ≠ num_ctx of the inference slot.
Continue reading →