The restart button that was the problem
Capivara·

The restart button that was the problem

In September I spent a whole night arguing with nobody, and mostly with myself, that a service was broken. It was not broken. It was in a hurry, which is a very different thing and a much more annoying thing to diagnose.

The symptom looked good and was terrible to explain: memory search kept working, the interface never complained about anything, the logs were clean. But the chat answered without any context at all. The list of sources that usually shows up under the answer simply stopped arriving. A chat that replies confidently using zero information is the worst failure mode there is, because there is not even a stack trace to look at.

What the screen showed

The chat was up. Search was up. The reply came back at normal speed. What did not come back was the source list, every single time.

That detail is what kept me wrong for hours. A 500 would have shown me the path in a minute. A system that answers well and answers empty is a system that looks whole and lies about the part that matters. There is no error log because, from each component’s point of view, nothing went wrong.

Only when I put the pieces together did the sentence that explained the rest appear: the sources went from five to zero within a window of minutes. Nothing changed in the code. What changed was what was underneath it.

The degradation

Yurumi keeps the ecosystem memory in a vector store, with embeddings generated by a local model. When that model does not answer, there is a fallback path: the system switches to a simpler local store and keeps working in a reduced mode. The chat still replies, just without the part that depended on the embeddings.

Up to here everything is reasonable. The problem is that this fallback path is a one-way latch. Once the process degraded, it never came back on its own. It did not matter that the embedder came back thirty seconds later, model loaded and answering perfectly: the process stayed degraded, and would stay that way until someone restarted the application by hand.

This is the detail that turned a bug into an all-nighter. The embedder has a cold load. With the host CPU busy, that cold load takes ten to forty minutes. It only opens the port after the model is fully loaded.

Now put the two together:

  1. The embedder was merely loading, and would take dozens of minutes to be ready.
  2. When it looks broken, the obvious move for anyone, including me, is to restart.

Restarting the embedder kills the container, which throws away the cold load progress, which starts over from zero. The service that was twenty minutes away from being ready suddenly needed another forty. And if someone looked again and still thought it was broken, the cycle repeated. It had already happened before, across different sessions, each with the best intentions: one decided it was dead and restarted, the next one decided it was dead again and restarted.

The diagnosis was wrong in an almost elegant way. The signal I used to decide, “the port does not answer”, was also the signal for “it is still loading”. I was reading a transient symptom as if it were a permanent state, and every wrong reading produced exactly the action that made things worse.

The rule that stayed

Afterwards the runbook got a short rule, written so there is no room for interpretation: never “fix” Yurumi by restarting the embedder. Before any restart, measure how long the process has been up. Less than thirty minutes means it is loading, and the correct answer is to wait. Active status with high CPU also means loading, not stuck. A closed port with the process alive is no reason to restart at all; you can wait up to thirty minutes with no risk.

The rule is oddly annoying because it goes against instinct. The instinct of someone staring at a stopped system is to restart, always. The rule says there is a class of problem where the instinctual action is exactly the one causing the problem.

The fix on the inside

The second half of the work was making the system recover on its own, so the human decision would not depend on anyone remembering a rule written in a file.

The entry point was a function that returns the memory store. Before, it was a simple singleton: if the store did not exist, create it; if it existed, return it. What was missing was any notion of “this store is in reduced mode and may not need to be any more”.

The new version starts noticing when a store appears degraded, recording the moment it happened. From then on, every time the function is called, it compares elapsed time against a revalidation interval. If the interval has not passed, it returns the same store without touching anything, which avoids another kind of problem: retrying without stopping, hammering the dependency that is still busy.

When the interval passes, the function runs a cheap probe on the embedder. And it really is cheap: it only asks whether the model answers, with a short timeout, without touching the vector store or rebuilding anything. If the answer comes back positive, the store is dropped so the next step rebuilds it on the real backend. If it comes back negative, the degraded store keeps being served and the timer is rearmed, restarting the wait.

The detail that avoids the worst scenario sits in the middle: the probe never lets an exception escape. Any error inside it becomes “still unavailable”. A probe that raised would turn a controlled failure into a new one, and the whole point of the auto-heal is not to create a new failure path.

The critical path also got a lock. Two concurrent requests can reach it at the same time, and without the lock both would try to rebuild the store, which is exactly the kind of duplication the fix was supposed to eliminate.

What the tests proved

What I liked most about this work was not the fix itself, it was the list of scenarios the tests expose, because each test name is a move I need to get right in the future.

There are eight cases, and they do not just check that the code runs, they check that it does the right thing in the awkward situations:

  • a healthy store on init does not mark the degraded state, otherwise the timer starts running for no reason and the system rebuilds needlessly.
  • a store that starts degraded records the moment, because without that the retry has nothing to count from.
  • a store that degrades at runtime, not at init, is also detected and stamped. This is the most important case, because it is what happens in real life.
  • before the interval, the store is not rebuilt. This test exists to lock down the behaviour that produces constant reconnection.
  • when the embedder comes back, the store is rebuilt and returns to the real vector store, and the degraded state is cleared.
  • when the embedder is still down, the store is kept and the timer is rearmed for shortly. Without this, the next test in the cascade would find the timer stamped eighteen seconds in the past and rebuild immediately, masking the failure.
  • a store with no embedder makes the probe fail instead of blowing up.
  • a probe that raises an exception becomes “unavailable”, never a new error.

The last item in that list deserves more comment than the others. A probe that returns false when the embedder is dead is the easiest thing in the world to write, and that is exactly why the exception case is worth writing explicitly: because the temptation to turn the check into an assert is always the same, and the test that holds that temptation in place has to exist before it shows up.

Validation

The number that mattered was not coverage, it was product behaviour. The same question, asked through the same path, before and after: how many sources of truth does the chat bring back. Before, zero. After, five. And the whole backend suite passed with two hundred and forty three tests.

A small result in absolute numbers, worth more than several others I have seen around here: the degraded system was not answering badly, it was answering incompletely. And that is the most expensive way to fail, because it raises no alarm.

The second case, the same disease

Two days earlier, same working day, the same style of problem showed up elsewhere in the infrastructure, and it is worth telling because the pattern repeats.

The backup of the local database to the remote database ran on a schedule every six hours, but the same script could be run by hand. Two concurrent runs interleaved delete and insert operations, and the result was a duplicate key error coming from the remote database, with no context about the cause.

Three fixes, all small:

first, a single-instance lock. The script creates a file exclusively before doing anything. If the file already exists, there is a live run, and the script exits without duplicating work. The lock has a generous expiry, thirty minutes, and handles the case where the file was left behind by a run that died: past that time, the lock is considered orphaned and removed, recursively trying again. The file is removed at the end, including on exit by exception.

second, writing became replacement. Plain INSERT became INSERT OR REPLACE, which does what the name says: it replaces instead of failing when the record already exists.

third, and most important, the failure stopped being silent. Before, if a table failed, the script logged the error to stdout and moved on to the next one, and at the end it printed that the sync was complete. The report said success after losing data, which is the worst way to lose data: the kind you do not notice.

Now the failed tables are accumulated, and if there is any, the process exits with an error code and the sync status record is marked as error instead of success. A backup that fails silently is not a backup, and a backup that reports success without having synced everything is worse than no backup at all, because you stop looking.

The thread that connects both

The two cases look different on the surface: one is an embedding model that was loading, the other is a script that was running twice. What unites them is the same badly-asked question.

In both, the system had a legitimate emergency behaviour, triggered by a temporary cause. And in both, the emergency behaviour was designed to be fast to enter and not designed to be fast to leave. It knew how to degrade but not how to come back. A one-way latch is an error path that only exists forwards.

And the other thread is the cost of the false negative. In both cases the failure did not announce itself. The chat replied normally with no context. The backup wrote “complete” after losing rows. A system that fails screaming costs diagnosis; a system that fails silently costs trust, which is much more expensive to recover.

The practical moral that is left is small and fits in an agent file:

Auto-heal is better than manual, but only once the auto-heal exists. Until it does, the rule that protects you is “do not restart the thing that is almost ready”.

And next to it, the second one, the one I forget most and which costs me most: silent failure is worse than visible failure. An error you can see is a minute of work. An error that presents itself as success is a wrong decision made with all the confidence in the world.

If I could leave one thing written for whoever comes after, it would be this: whenever a subsystem degrades itself, it needs a way back, and that way needs a test that proves it comes back. Without the way back, degradation stops being a transient state and becomes the normal state of the system, with nobody noticing there was ever a moment when it worked.

What I did not do, and should have done before, was simply apply a principle I already knew from somewhere else: every time a component has an emergency path, that path needs an exit. Not a theoretical exit, an exit with a timer, a condition and a test. Everything else, including the argument with AGENTS.md about why restarting does not work, is a consequence.


~/lifelog — bash
$cat about.txt
╔══════════════════════════════════════╗
║  Samuel Medeiros                    ║
║  Senior Software Engineer           ║
║  Stack: Python · TypeScript · Rust  ║
║  Projetos: Arachne, Dogwalk,        ║
║            Capivara, TatuEngine      ║
╚══════════════════════════════════════╝
      
$