One person's key, another person's bill
Arachne·

One person's key, another person's bill

Arachne has a feature I always thought was honest: everyone brings their own model key and pays their own bill. No accidentally burning someone else’s quota, no company key mixed into the middle. The entire system depended on one simple thing to keep that honesty — that one person’s key was kept for one person only.

That simple thing was a field. And the field was optional.

The mental model was right

Internally Arachne has what I call a single point of model choice. A resolve function follows a fixed, explicit order: first try the caller’s own key, then try the workspace key, then the platform key, and finally the local model. The reason is recorded alongside it, in a field that says where the payment came from — because the key’s cost belongs to its owner, not to us.

That order is written in the code as a cascade of four if statements, each with a comment explaining why. Reading it is comforting: you can see the intent of whoever wrote it, you can understand why each step exists. The problem was never the order. The problem was the second step.

The second step is the legacy path: configs that existed before there was such a thing as an owner. And a legacy config, in practice, means a row in the database with no owner filled in.

The line that decided everything

The legacy path filters the workspace configs looking for one that fits. And the choice, in the old version, was a single line:

cfg = None
for c in rows:
    cfg = c
    break

First active config in the workspace, without asking whose it is. It reads as innocent compatibility code — you grab the workspace config because that’s what the old code did. It worked, because in the old model there was no such thing as someone’s key: the key belonged to the workspace, period. Everyone in the workspace used the same key because the key was the workspace’s.

Except the model changed. And the compatibility code stayed.

When configs started carrying an owner, two kinds existed at once in the same workspace: the new ones, with an owner filled in, and the old ones, without. And the compatibility line did not distinguish. It grabbed the first active config it found — whoever’s it was.

The detail that closes the trap: the filter did not even look at the owner. It wasn’t “grab the config with no owner to keep backward compatibility”. It was literally “grab the first”. If, in a shared workspace, the config with an owner had a smaller id than the legacy config — which is the normal case, because the legacy one was created first — the filter returned the colleague’s key, owner and all.

Why the bug only showed up in tests

The test suite for this wave has fifteen cases, and two of them are the heart of the story. One creates a config with an owner in the workspace and asks the legacy path what it returns when the requester is someone else from the same workspace. The expected answer is none — the filter must refuse. The other confirms compatibility still works: an old config with no owner must still be served by workspace, because nobody should lose their key because of a security fix.

The second test is what stops the fix from becoming a silent regression.

What those two tests document together, and what stops me here, is the detail: this was one of the few security bugs I’ve seen written as two behaviour tests rather than a list of “must not do this”. The test says what the system does when the scenario is hostile. “When a config has an owner and the requester is someone else, the legacy path returns nothing.” “When a config has no owner, the legacy path still returns it.” Two sentences that, together, describe the entire boundary. The rest — the line of code, the filter’s name, the signature — is implementation. The sentences are the contract.

The fix

The filter gained the condition it was missing. A config with no owner still stands — compatibility. A config with an owner is only served to its owner. And that last part is what truly closes the hole: it’s not just “no owner passes”. It’s “someone else’s owner does not pass”. Any key with an explicit owner stays invisible to anyone who isn’t that owner, even living in the same workspace. That holds even for the legacy path: a shared workspace stopped being a back door.

cfg = None
for c in rows:
    if not c.owner_email or c.matches_owner(email):
        cfg = c
        break
if not cfg:
    return None

One extra condition. matches_owner normalizes before comparing, because comparing emails with different case and spacing is the kind of detail that makes a fix look like it doesn’t work: the owner is “Owner@example.com” in one place, “ owner@example.com “ in the request, and a plain == returns false — and then the system “protects” the key from its own owner. An ownership guard that gets the owner wrong is worse than no guard at all, because it looks like it’s working. Normalizing is what makes the guard verify who the person is rather than how they spelled their address.

The part I almost forgot

When the fix was ready, the detail that made me stop was the key-test endpoint. It didn’t just guard the config — it fired the key. It was a meter: it sent a real request to the provider using someone’s config. Guarding the get was half the work; the /test was a button that, with the right id, measured the colleague’s key. What caught it wasn’t code reading — it was re-reading the whole route asking “what does this route do, not what it looks like it does”.

That’s what made me understand that ownership must be checked on every operation that touches the key, not just on reads. Get, update, delete, test — four verbs, one rule. Having the ownership guard in the right place is what makes “not yours” become the same answer across all of them. And the right answer to “not yours” is 404, not 403: 403 confirms the resource exists. Whoever receives a 403 knows the id exists and that it belongs to someone else — which is already information. 404 lies about nothing.

What I take away

Two things, and neither is about security.

The first: an optional field that expresses ownership is a trap. If a field carries “whose is it”, it should be mandatory the moment the concept appears. owner_email was born optional because legacy configs existed, and so compatibility became an if — but that if did not distinguish “has no owner” from “has a different owner”. When the rule is “no owner passes, someone else’s owner does not pass”, optionality is no longer a migration detail: it is the bug’s own surface.

The second, more expensive: a legacy path is a place where the new rule never arrives. The new rule — the key has an owner — was written in the new path, and the old path kept the old rule. A compatibility path isn’t just code that still runs; it’s code that still decides. And as long as it decides, it decides wrong.

Neither fix required a revolutionary line. It required me to stop reading the filter as “compatibility” and start reading it as “authorization”. It wasn’t a filter. It was a badly written authorization, waiting for someone to stop and rewrite it instead of accepting what was already there.

~/lifelog — bash
$cat about.txt
╔══════════════════════════════════════╗
║  Samuel Medeiros                    ║
║  Senior Software Engineer           ║
║  Stack: Python · TypeScript · Rust  ║
║  Projetos: Arachne, Dogwalk,        ║
║            Capivara, TatuEngine      ║
╚══════════════════════════════════════╝
      
$