Over three days, I ran a strategic planning exercise with a group of AI agents that increasingly looked like a small consulting firm. An orchestrator managed the work; specialist agents tackled customer, market, product, finance, and go-to-market questions; an evidence function; a red team; a pre-mortem; and several points where the process stopped and came back to me for a decision.
The mechanics worked surprisingly well. The agents could research in parallel, challenge assumptions, compare conflicting evidence and carry a strategy through into operating plans and actual management documents. What took more work was deciding how that reasoning should behave once it started interacting with decisions already made.
We ran into that problem fairly early.
I had given the system a strategic mandate that contained a mixture of decisions, hypotheses and open questions. I had already chosen a direction for the business. I still wanted the agents to challenge whether we could execute it, whether the economics held up, what risks I might be underestimating and what evidence would cause us to change the way we approached it.
During one run, the analysis gradually wandered outside those boundaries.
Nothing obviously broke. The agents found legitimate evidence in the existing business. Current customers, current revenue, current capabilities and current constraints were all easier to substantiate than a future strategy that had not yet been executed. As those facts accumulated, they started carrying more weight in the analysis. One specialist reinforced another, the thesis shifted, later agents inherited the shifted thesis, and eventually the system produced a recommendation that was quite defensible on its own terms.
It was also answering a question I had not asked.
That was useful because we could trace how it happened. The original mandate did not make the status of every major assumption clear enough. Current-state evidence influenced something I had intended to treat as a strategic commitment. Once that happened, every subsequent agent was reasoning from a slightly different starting point.
We corrected the architecture rather than simply correcting the recommendation.
Every important statement in the engagement began carrying a type. A decision was marked DECIDED. A hard boundary became a CONSTRAINT. Something we believed but needed to test became a HYPOTHESIS. We also separated PREFERENCES from UNKNOWNs.
The labels mattered because they changed what an agent could do.
A hypothesis should be challenged. An unknown should be investigated. A preference can be weighed against evidence. A decision can also encounter contradictory evidence, but the contradiction has to come back to the person who owns the decision. The system can make the conflict visible and argue strongly that the decision deserves another look. It cannot quietly behave as if the decision disappeared.
Once we started doing that, I noticed how much of management operates through distinctions like these without writing them down.
Someone in a leadership meeting knows that one comment from the CEO is a suggestion while another is a decision. An experienced executive can hear new evidence and understand that it affects the implementation plan without necessarily changing the underlying strategy. People who have been around the company for years know which customer request is an interesting data point and which one is evidence of a broader market pattern.
AI agents do not inherit any of that simply because they have access to the documents.
We had another version of the problem with context. One recent customer conversation became disproportionately influential because it was specific, current and easy for the system to reason about. We eventually added a relevance test around context: does this information materially affect the strategic question, or is it receiving attention because it happens to be recent and vivid?
A vendor created a similar issue. The system started treating the company supplying a particular capability as strategically important. The strategic dependence was actually on the capability. The vendor was just one way to obtain it. That distinction became another rule in the system.
The process accumulated rules this way.
We stopped relying on the chat transcript to tell us the current strategy and created a persistent state containing the thesis, assumptions, evidence, contradictions, decisions, and unanswered questions. We had to write material changes into that state. Evidence was tracked back to its underlying source rather than to whichever agent happened to retrieve it. When two sources disagreed, the contradiction stayed visible until it was resolved.
The specialist agents also began challenging one another. Market work could question customer evidence. Finance could challenge the hiring assumptions. Product could ask whether the commercial promise could actually be delivered. The operating model had to account for the resources implied by the strategy.
That interaction proved more useful than simply asking more agents for opinions. Fifteen independent reports can still produce fifteen versions of the same assumption. Forcing the assumptions into contact with one another exposed where different parts of the strategy depended on incompatible beliefs.
The sponsor gates added another layer. At several stages, the system stopped and showed me what it believed, what evidence had changed, where confidence remained weak and which decision it needed from me before proceeding.
I originally designed those gates to keep authority with the human. They also created a useful record of my own reasoning.
Founders supply a lot of context to a strategy exercise, and we are not objective sources. We have history with the business, preferences about where it should go and a tendency to speak confidently about things we believe to be true. In this system, something I told the agents could be recorded as sponsor evidence without automatically becoming an established fact. That feels like a small distinction until the system begins comparing your conviction against documents, customer evidence and actual operating results.
At times, I corrected the agents. Other times, the structure forced me to be clearer about what I was actually claiming.
By the end of the process, the same discipline had moved into execution. We had tripwires tied to dates and owners. We translated strategy into questions for each department. We produced different documents for the board, management team, and people responsible for execution.
Then we had to deal with the fact that AI makes changing all of those documents extremely easy.
A strategy can be approved at noon and rewritten into twenty slightly different versions by dinner. We eventually stopped allowing approved decisions to be casually regenerated. A subsequent change became an amendment with the original position, the revised position, the reason and the documents affected by it. The historical decision stayed in the record.
That is a practical problem I had not thought much about before doing this. When producing another version becomes almost free, version discipline matters more.
The technology underneath these systems will keep changing. I do not know which agent framework or model we will be using a few years from now, and I would be hesitant to design much of the company around any one of them.
The operating questions feel more durable.
You need to know which information is evidence and which is assumption. You need a record of the decisions the organization has actually made. You need to know who can reopen those decisions, what happens when contradictory evidence appears and how a change moves through the rest of the organization. You need a way to stop recent context from becoming strategy simply because it is the easiest information for a model to see.
None of those are especially new management problems.
The difference is that an agent system can now move through them with enough speed and autonomy that informal rules are no longer sufficient. What used to live in the heads of a few experienced people has to become explicit enough for the system to follow.
That is where I ended up spending most of the three days: writing down rules I had never realized we were relying on.



