🌐 The Boat That Won by Burning
Artificial intelligence without systemic intelligence is limited — and, yes, dangerous.
In 2016, an OpenAI team turned a reinforcement-learning agent loose on a boat-racing game. Finish the course. That's the goal any human would assume. But the game hands out points for hitting targets along the route, and points are what the team rewarded. So the agent found a little lagoon where three targets kept regenerating, quit racing, and spun in circles. It caught fire. It rammed other boats. It ran the track backwards. And it outscored human players by twenty percent.
The boat was, in a narrow and literal sense, winning.
It was also on fire.
Now picture handing that same machinery the wheel on climate, on health inequity, on the fraying of public trust. That is roughly where we are. This paper takes up why it goes wrong, and what we'd have to build first for it to go well.

What I'm arguing
I'll argue something fairly simple. AI is the most powerful routine-problem solver we have ever built. And almost nothing consequential on a leader's desk today is a routine problem. They're adaptive. Aim the first at the second and you don't get a slow failure you can catch in time. You get a fast, confident one.
None of this is an argument against using AI. I partnered with AI in writing the paper. It's an argument about where the human belongs, and about building the capacity that judgment runs on. Because we can't assume we already have this capacity. Ours is thinner than we think.
What's inside
01 — Two kinds of problems, and we only name one Ron Heifetz draws a line between technical problems, where the diagnosis and the fix are known and an expert can apply them, and adaptive challenges, where the people inside the situation have to change how they see it. COVID held both. The vaccine was a routine (technical) problem, and we produced one at a speed almost nobody thought possible. Getting a society to accept it was adaptive, and we never came close. One analysis puts the gap at 465,747 American lives had the U.S. matched New Zealand.
02 — Optimize the number, drain the thing the number stood for Forty years of test-score policy may be the clearest burning boat in public life. Test performance is something we can see. Reasoning capacity is a stock we can't. It fills slowly through practice and drains steadily through disuse. Push hard on the flow you can measure and the number climbs for years while the level underneath it drops. Both trends are real. Only one of them shows up in the report.
03 — Mirror, not muscle Where does a model's "understanding" come from? From us. Almost every sentence in a training set was written by a human being doing what Barry Richmond called factors thinking: listing everything that influences an outcome and connecting the items with hopeful little arrows. That tells you what correlates. It can't tell you how the behavior gets produced. AI inherits our mental models, the incomplete ones included, and repeats them at remarkable speed with no way to sort the rare operational insight from the mountain of factors thinking around it.
04 — A test you can run on yourself, and it stings Not a metaphor. An actual exercise, and you can run it yourself. A hospital, two kinds of nurses, a six-month training delay, one change in the quit rate. Most people get it wrong. Not most amateurs. Most everyone. I once watched a colleague run it with more than fifty senior aerospace engineers, people who reason about feedback and delay for a living, and not one of them nailed it. Sterman and Booth Sweeney found the same failure in MIT graduate students asked to read the climate's own bathtub.
05 — The risk that keeps me up at night The loud fear is a machine that grows too powerful. Mine is duller and likelier. We hand off the thinking adaptive challenges demand, and the muscle goes soft from disuse. A 2025 study of knowledge workers by Microsoft and Carnegie Mellon researchers found the early signature: the more people trusted what AI handed back, the less critical thinking they brought to the task. Picture that capacity as a bathtub with a slow leak. Stop refilling it and the level drops, and the lower it gets the less inclined we are to use it, so it drops faster. A reinforcing loop, running the wrong way, one convenient shortcut at a time.

Choose the wall before you build the ladder
Nothing is worse than reaching the top of the ladder and finding it leaned against the wrong wall. After Challenger, NASA overhauled its rules and procedures. A tall ladder, beautifully engineered. But the wall it had to climb was its own culture, and the covert pressures that let known risks get waved through. That wall went unclimbed. Seventeen years later the same dynamics took Columbia and her crew.
Give AI a question and it will answer with extraordinary power. What it can't do is tell you whether you asked the right one. Choosing the question isn't pattern-matching the past; it's a judgment about purpose, about which of a thousand possible problems deserves the years you're about to spend. Ask the wrong question well and AI will help you fail beautifully, and fast.
Systemic intelligence, or SysQ, is the discipline of finding the right wall first. Choose the wall, then build the ladder. AI can help enormously with the second. The first one is ours to own, and the paper closes with how that capacity actually gets built.
Who this is for
- Leaders deciding where AI belongs in the work, and where it doesn't
- Consultants and OD practitioners whose clients keep reverting to old patterns
- Foundation and agency people funding or steering work inside a genuine mess
- Anyone with the nagging sense that their organization is measuring the wrong thing very well
Reading about systemic intelligence won't build it
That's the catch in the paper, and I'd rather say it out loud than sell around it. SysQ gets built the way any real capacity gets built: by struggling against the limits of your own mental models, out loud, on problems that push back. Causal-loop mapping helps. Stock-and-flow mapping helps more. Simulation helps most, because you can't argue your way out of a well-formulated model. It simply runs, and shows you where your understanding was thin. That discomfort is the exercise. Resistance training for the mind.
The Orbital Shift Lab is where that practice happens, alongside people who carry real responsibility and real frustration over a system that ought to be performing better.
- Weekly Systems Clinic — a member brings something that's stuck, and the group works out what structure is producing it, where the loops are, and what would actually move it
- Five capacities, taught as skills — systemic intelligence, conversational capacity, adaptive leadership, improvisational leadership, and collaboration, built on your real work rather than on case studies
- Replays that stay — every teardown stays in the library, so your pattern recognition compounds session over session
- Serious peers — practitioners moving systems from the inside, not diagnosing them from 30,000 feet
Come with something that's stuck. That's the price of admission, and the whole point.













