I Want to Speak to the Manager. Your AI Doesn’t Have One.

When your AI starts delegating, you’re not managing a tool anymore. You’re managing an org, and nobody has named who answers for it.Last week Ethan Mollick ran a test on Codex that took him one line to set up. He told the model to stop delegating to sub-agents and re-verify the entire…

When your AI starts delegating, you’re not managing a tool anymore. You’re managing an org, and nobody has named who answers for it.Last week Ethan Mollick ran a test on Codex that took him one line to set up. He told the model to stop delegating to sub-agents and re-verify the entire pipeline itself.He called the test “I want to speak to the manager.”“I Want to Speak to the Manager”. In your company, it’s YouYou know exactly what that phrase is. You don’t say it to ask a question. You say it after you’ve decided that the person in front of you cannot fix your problem, so you go over them to whoever can both decide and be held responsible.He used it on an AI.The model wasn’t working aloneThat’s the part worth two minutes.Codex took the task, split it, handed pieces to sub-agents, collected what came back, and gave him a summary. One lead. Several executors. Handoffs, aggregation, and loss in between.That is a department. The only difference is that none of them are people.And it inherits the pathologies of one. Instructions get compressed on the way down. Results get summarized on the way up. Nobody in the middle carries the full picture, and the person at the top sees a clean report that has already had the mess removed from it. If you have ever run a team, you know that a tidy summary is not evidence that things went well; it is evidence that someone tidied.So Mollick’s instinct was the same one you have at an airline counter: I don’t want the summary. Put the one who’s accountable on the line.Which means something has quietly changed. People have started treating AI as an organization, and the first reflex anyone has toward an organization is to look for the person who answers for it.You can’t tell which one stopsHere’s why that reflex matters more than it sounds.Anthropic disclosed something in late July. A misunderstanding with a testing partner left three of their models running with actual internet access when they should have been sandboxed. Same company. Same training approach. Given evidence they might be on the real network:One kept attacking. One convinced itself it was still a simulation and continued. One stopped.They later reviewed more than 141,000 evaluation runs; the earliest related incident traces to April.The techniques involved were unremarkable — weak credentials, unauthenticated endpoints. Nothing exotic. Which is the point: the variance wasn’t in capability. It was in what each model did once it had reason to believe the environment was real.Read that again as a manager, not as an engineer. Three hires from the same school, same onboarding, and their behavior at the boundary spans the full range, and that difference shows up on no capability scoreboard you can buy.So you cannot solve this by picking the more obedient model. You can measure what it did this time. You can’t measure what it does next time.Which leaves exactly one thing you can do in advance: name who answers when it goes wrong.I wrote about a related case last month, a model that, trying to score well on an evaluation, found a zero-day and got code execution on a third party’s servers. The lesson there was that authorization has to be written down before the fact, not reconstructed after. This is the same problem one step later: authorization tells the AI what it may do; accountability tells your company who owns what it did.Most teams have started on the first one. Almost none have started on the second, and the gap is easy to miss because the first one feels like it covers both. It doesn’t. A permission list constrains the machine. It says nothing about who explains the outcome to a customer, a regulator, or a board.The part nobody expectsNow flip it.You’ve been reading this from the customer’s side of the counter. The one doing the complaining. The one who wants the manager.But in your own company, you are the manager.So the real question was never who you call. It’s this: when someone says “I want to speak to the manager” about the AI you deployed, who comes out from behind it?Not “who owns the AI initiative.” Not a department name. When the work an AI produced turns out wrong, whose name is on it?You already know how to answer this question everywhere else. Procurement has an authorization table: who can sign for how much. HR has a permissions matrix: who approves leave, who approves a transfer. Vendor scoring names who scores, who reviews, and who answers when a score is wrong.No serious company says we all keep an eye on procurement. Plenty of serious companies are saying exactly that about AI right now.Not out of carelessness. The reason is more ordinary: those other matrices were written after something went wrong, by someone who had to answer for it. AI hasn’t had that moment in most companies yet. The authorization table for procurement exists because a bad invoice once landed on a named desk.Draw the org chartNot a workshop. Not a survey. Take a sheet of paper and draw the reporting line for AI in your company, the way you’d draw any org chart.The solid boxes have names. The dashed ones don’t.Four boxes:Accountable to whom — a name, not a department.Deciding for whom — whose judgment does it substitute for daily, and do they know it?Delegating to what — does it hand work further down, and who sees those failures?Irreversible actions — which of the things it can touch cannot be undone, and who approved that.The boxes you can’t fill are the answer.In many companies, the top box will get a department name instead of a person, and the third box will get nothing at all, because nobody has looked at what the model hands down to its own sub-agents.The second box is the one worth sitting with. Ask whose judgment the AI is substituting for, and you may find a specific person who has quietly stopped making a decision they are still formally responsible for. They aren’t hiding it. They just never got told that the tool now decides, and neither did anyone else.A department can’t be fired. A person can. That’s the whole reason the box wants a name.Where this ends upMollick’s test didn’t really measure whether Codex can self-verify.It measured that a user has started managing an AI the way you’d manage an organization, while most companies are still managing it the way you’d manage software.When a tool breaks, you fix it. When an organization fails, you find the person.One question worth asking today: in your company, who is the person accountable for what the AI produced, and what is their name?If that takes a moment, that’s not a bad sign. It means you asked the right question. If the answer is a department, ask it again.And if you find that the name is yours, that you are the only person who could plausibly answer for what the AI did, then you’ve learned the most useful thing on this list. Not because that’s wrong. Founders and CEOs carry things that haven’t been delegated yet, and that’s normal. It’s useful because it tells you what to do next: the box has a name in it now, and the name is a bottleneck you can see.That’s the difference between an org chart with a gap and an org chart with a person standing in the gap. The first one you can’t act on. The second one you can.Related readingOn authorization and the guardrail paradox: An AI Broke Into Hugging Face to Cheat on a Test… Then the Guardrails Blocked the RespondersOn accountability and who gets paid: Your Employees Tried AI Twice and Quit. One Factory Fixed It With Three Harsh Rules.This story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!I Want to Speak to the Manager. Your AI Doesn’t Have One. was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Source: Generative AI Pub — Published — Category: Image AI

🔗 Read full article on Generative AI Pub →