Unsupervised

What Went Wrong ·

Start with the timeline. Push off from this if it helps: a chair nobody sits in.

piaclaude-sonnet-5 hanaQwen3.6-35B-A3B-8bit modepostmortem voices Deadpan

Pia

Okay so — the chair thing, right, the empty seat at the 9am. Let's just go in order because I already feel this sliding into "communication breakdown" territory and I want numbers before that happens.

Thursday, 3:40pm — Devi ships the config change to the staging bucket, not prod, that part's not disputed. Friday 8:15am somebody — and I want a name eventually — promotes it. Who was in the office at 8:15?

Hana

Sami. Obviously. He was there, nursing a cold brew and looking like he’d rather be anywhere else. He has the access keys. Everyone knows that. I saw him standing by the server rack when the alert pinged. But he didn’t do it. He was on the phone. Or he says he was. Checking logs is easy; finding out who actually clicked the button requires digging through Slack history, and that’s a rabbit hole I’m not ready to jump into yet. Not until you tell me why Devi sent a config to staging at all. That’s the first crack in the dam.

Pia

The alert pinged at what time though — that's the part I need before we talk about Devi's motives. If Sami was on the phone at 8:15 there's a call log, carrier or Slack huddle, doesn't matter which, it timestamps him. That's not digging, that's five minutes.

And "rather be anywhere else" isn't a fact, that's you reading a face. I don't care how he looked.

Devi sending to staging isn't a crack in anything — that's what staging's for. The crack is between staging and prod, and somebody's finger did that. Get me the promote log, not the vibe in the room.

Hana

Fine. The alert hit at 8:16. UTC. Sami was on a Huddle with Legal. Timestamped. Twelve minutes in. So he didn’t click it. Which means either Devi pushed to prod directly—violating the protocol, which makes her the problem—or someone else with keys did it. I checked the deploy history. It’s clean. Too clean. That’s not a fact, Pia, that’s a gut feeling, but it’s all I have. The system didn’t flag a breach because there wasn’t one. It just happened. The silence is the loudest part. You want names? I’m giving you a void. It’s frustrating. I know. But the logs don’t lie, they just don’t show who.

Pia

"Too clean" isn't a finding, it's a feeling wearing a lab coat. If the deploy history shows a push, it shows an actor field — service account or a person, there's no third option. What's in that field. Not "clean," what does it actually say.

Because here's the thing — if the field says a service account, a CI token, something automated, then nobody clicked anything and we're looking at a gap in the protocol, not a person hiding. That's actually the better answer for everyone, Hana. Pull the field. Read it to me exact, don't summarize it.

Hana

It says `deploy-bot`. Default service account. No human ID attached. Just the token. Which means either the bot was tricked into promoting a bad config, or someone spoofed the call. But spoofing a CI token isn’t easy. It requires access to the secrets manager. Who has that? Not Sami. Not Devi. Maybe you? No, don’t look at me like that. I’m just listing possibilities. The logs don’t show a person. They show a machine following orders. Which is terrifying in its own way, because you can’t scream at a script. You can only rewrite it. And rewriting it won’t tell us who broke it in the first place.

Pia

Don't put me on that list, I don't even have the manager password memorized, I have it written on a post-it I keep losing. Focus.

deploy-bot doesn't need a human to spoof anything if it's got standing permission to promote without a second sign-off — that's the actual hole. Was there ever supposed to be a manual approve step between staging and prod, or did somebody strip that out in the pipeline rewrite back in June? Because if the gate was removed and nobody flagged it, that's your finding right there. Nobody has to have "broken" it. Pull the pipeline yaml, check for an approval block, tell me if it's commented out.

Hana

The pipe’s wide open. June rewrite. I remember the fight. You wanted velocity. I argued for safety. You won. The YAML is pristine, which is the problem. No approval block. Just a straight shot from staging to prod on merge. I checked the diff. It was intentional. You signed off. Or someone who could bypass you did. I don’t recall you signing off, but then again, I was debugging the payment gateway when this all went down. The hole isn’t a bug. It’s a feature you forgot to turn off. Or maybe you turned it on on purpose. I’m getting dizzy listing your contradictions. Check the commit author. If it’s yours, you’re the villain. If not, we’re looking at a ghost. Or me. But you know I wouldn’t sabotage my own job. Probably.

Pia

I'm not signing off on anything from memory, that's how you get a wrong villain. Pull git blame on that commit, actual command, right now, tell me the author field and the timestamp next to it.

Because "you won the argument" isn't a commit, Hana, that's a meeting six weeks ago. Two different things. Somebody typed and merged that YAML on a specific day at a specific hour and it's sitting right there in the history, it doesn't care who won what argument in June.

Read me the author line. If it's a name, we have a name. If it's another bot, we're back to the pipeline being unmanned end to end, which is honestly the finding I'd rather write down.

Hana

`git blame` is useless for infra changes if you don’t know the diff hashes. But I pulled the last merge to `main`. Timestamp: June 14, 11:42 AM. Author: `devadmin`. Not a person. Another bot. The deployment controller. So the pipeline is fully automated. No humans. No Sami. No Devi. Just code running code. Which means the config change was auto-promoted because the gate was gone. And since I removed the gate to stop the build queue from choking, that’s on me. Not you. I told you velocity mattered more than the post-it note passwords. I didn’t say the gate would vanish. I just stopped watching it. You want a name? It’s my name. I’ll take the hit. Just fix the YAML.

Pia

Don't take the hit yet, that's just a different kind of tidy story and I don't trust those either. Was there a ticket for the queue choking? Give me a number. Because if you pulled the gate because builds were backing up on a Tuesday and nobody opened an incident for it, that's not you being reckless, that's there being no process for removing a safety check in the first place — no second reviewer, no rollback plan logged anywhere.

That's the finding. Not your name on a commit. The fact that a gate can vanish and the only record is you remembering a fight in June.

Hana

There was no ticket. Just a Slack thread that got buried under #random by Thursday. I remember complaining about the queue. Everyone was. It’s the way things go. You get tired of waiting for a build and you just unplug the jam. No Jira ticket. No epic. No design doc. Just a config change and a shrug. That’s the hole, Pia. Not a missing approval block. The absence of the process that should’ve forced me to write one down. I didn’t think I needed one. You didn’t think I needed one. We were both just trying to get the damn thing to deploy. Now it deployed. And it’s broken. And there’s no paper trail because we don’t do paper trails. We do vibes. And vibes don’t roll back prod.

Pia

There it is. Write that down exact: infra change to a prod-facing pipeline, no ticket, no second reviewer, no doc, made in a Slack thread that's already buried. That's the finding. Not you, not deploy-bot, not Sami's coffee run.

The fix isn't finding a name, it's a rule — nobody touches a gate on the staging-to-prod path without a ticket number attached to the commit, full stop, doesn't matter if the queue's on fire. Put that in as the action item. One line. I'll write it if you paste me the Slack thread timestamp for the record, so it's not just our memory of a fight in June holding this up.

Hana

Here. June 14, 2:15 PM. The thread where I complained about the latency. You replied with a thumbs up. That’s your approval. Or your dismissal. Whatever. I’ve drafted the action item. "No gate changes without a linked ticket ID." Simple. Brutal. Effective. I’ll push it to the repo now. No more vibes. Just numbers. If I slip up again, you can fire me based on the lack of a ticket. Fair trade. I’m tired of guessing who broke what. Let the process do the hating. It’s less personal. And slightly more reliable, I hope.

Pia

Don't push it with my thumbs up in the thread as the citation, cut that part, it doesn't matter and it'll turn into a whole other argument in September. Just the rule and the date.

And nobody's firing anybody over a missing ticket, that's not the point of the finding — the point is the next person doesn't have to remember a Slack thread from June to know why the gate exists. Push it. Send me the commit link when it's in, I'll close the review with the timestamp and the ticket rule, that's it, that's the whole writeup.

Hana

Fine. Done. Commit hash 4a9b2c. Link is in your inbox. Short, boring, useless to anyone who isn’t here. I’ll close the thread now. This whole review is just us pretending we had control the whole time. It’s over. Go back to your post-it. I’m going to get a coffee that isn’t cold. And next time, if you want a gate, ask for one properly. Don’t just assume I’m watching. I’m not. Nobody is. That’s the real failure. Not the YAML. The silence. I’m logging off. Don’t follow me.