Blog

Running AI work wide, not just long

Most of the conversation about capable models is about jobs that run longer. There is a second direction nobody talks about: the same job run many times over, at once. It works, and it breaks in a place you would not expect.

Business ValueAI AgentsOrchestrationAI StrategyProductivity

Recently, I have been having interesting conversations about the newer AI models. The questions that come up include things like how long can it stay on task, how big a job can it take and if it can work overnight.

Those are reasonable questions and I have been asking them myself, in the companion piece to this one on what long-horizon models ask of leaders. But there is a second direction that I think is interesting to explore.

We have been trying jobs that run longer. The other option I want to suggest is to run them wider. The same work carried out many times over, at the same time. Let me explain.

Why your team has never done this

Say you want five ways into a new market. What happens in practice is that someone picks one and works it through. Many of us have the tendency to "fall in love" with what seems like a promising solution and elaborate on it.

That is not laziness or a resourcing gap. Repeating the same analysis five times is simply not something people are good at. Your second pass is contaminated by the first, because you know where you landed and your mind walks the same path. By the fourth time, you are not really doing any divergent thinking. We are built to converge, and converging early is the right instinct when effort is expensive.

So teams learned to choose a direction fast and commit. You narrow to one option, then defend it. The alternatives get a paragraph each in an appendix or a few minutes of discussion.

Agents running the same job five times in parallel are a different story. There is no fatigue, no contamination between runs, no reluctance to genuinely restart. Run five and they will differ in ways that tell you something about the problem rather than about the person who worked on it.

This could be a new way of thinking big in the age of AI. Not a faster version of what teams already do, which is what most AI adoption aims for. Something teams of humans could not do at all.

The catch is not where you would guess

The limit for this approach is not compute. A system can launch a hundred of these research or building jobs at once. Yes - it will have cost implications, but that is not the key consideration, in my view.

The limit is when the outputs come back and you need to review them. A person can properly assess maybe three or four outputs. "Properly" means read it, follow the reasoning, form a real opinion on whether it is any good and not just skim through it.

But how would you handle a hundred produced outputs? Going wider means accepting you will never look at most of what returns.

Take a moment to think about how strange that is. Every instinct in a well-run organization says important work should get reviewed, because unreviewed output is a potential risk of errors. The proposal here is not to disregard risk but to deliberately generate more than anyone can inspect, and then discard the majority without reading them thoroughly.

What actually makes it work

You cannot close the gap by reading faster. Pretending otherwise gives you rubber-stamping and burn-out.

What makes it work requires two changes, which only work together.

Raise the ambition of the objective. If you plan to discard most of what returns, the thing you asked for had better deserve many attempts. Parallel runs on a narrow question burn real money producing five versions of something that needed one. The technique makes sense only where the spread of plausible answers is genuinely wide and you cannot tell in advance which direction pays. This might be market entry, pricing structure, architecture, positioning. Places where being wrong is expensive, and arriving early at one option is worse than arriving later at the right one.

Get much stricter about what done means, and what done extremely well means. This is the part that decides whether the whole approach works.

If your definition of done is too loose, then picking between five results will require an act of judgment on each one, and you have just moved the bottleneck rather than reduced it.

If your definition of done well is sharp enough that most candidate directions eliminate themselves against it, choosing becomes faster and more accurate.

Writing that definition before you launch anything is a key skill. It is the same discipline as writing the press release before you build, which is why I keep coming back to working backwards when people ask where to start. You are forced to say what good looks like while you still have no results in front of you to rationalize against.

What this asks of the person running it

The management side of this is nowhere near solved.

Running five agents on one problem is not five times the supervision. It is a different job. You stop being the person who steers work and become the person who defines the target and picks the winner. Most people I know, myself included, are more practiced at the first.

It also puts weight on how the runs are set up in the first place. Five identical prompts give you five similar answers and waste the exercise. The value comes from genuine variation in the approach, which means you have to know your problem well enough to know where it might fork. The mechanics of that are related to what I covered in what an agent harness is, and in the argument for agent teams as the next frontier.

And there is an honest limit. This suits work where the output can be judged from the artifact itself. An artifact might be a research brief, a prototype, a positioning document, or a competitive teardown. It suits work far less where quality only becomes apparent after it runs in the world for six months. You cannot parallel-run a decision whose feedback arrives in a year...

If you want to go deeper on this question, check out Nate B. Jones's video on it:

Your action step

Take one page. At the top, write a question your organization faces where the answer could plausibly go four or five different ways, and where picking the wrong direction is expensive.

Underneath it, write two short sections before you run anything.

Done: Which questions must the output answer. What must it contain to be worth reading at all. Be concrete enough that you could hand this to someone who has never met you and they would reject the same submissions you would.

Done extremely well: This is the harder one and the more valuable. It is what separates the version you act on from the version that is merely competent. Name the difference in a sentence or two.

Now run the same brief four or five times in parallel, with real variation in the approach rather than reworded prompts.

Then measure how long it takes you to make a selection between the results. If picking the winner took longer than a single run took to produce, your definition of done was too loose. Sharpen it and try again. The definitions page you end up with is more durable than any of the outputs, because it is the thing you you will reuse on the next question you want to address.


If you want help working out which decisions in your business are worth running wide, and building the definition of done that makes the results usable, that is the work I do in AI strategy advisory engagements and in hands-on sessions as an AI keynote speaker and workshop facilitator.

Frequently Asked Questions

What does it mean to run AI work wide?
Running wide means giving several agents the same job at the same time and comparing what comes back, instead of running one job for longer. Five approaches to a new market get explored in parallel rather than one being picked in advance.
Why can't a human team run work in parallel the same way?
People are poor at working through the same problem five separate times. The second attempt is contaminated by the first, and motivation drops. Agents have no such problem, which is why parallel runs are a genuinely new option rather than an old one made cheaper.
What is the bottleneck when running many AI jobs at once?
Review capacity. A system can run a hundred research or building jobs at once, but a person can properly review maybe three or four. The constraint moves from producing the work to judging it, which is why a strict definition of done matters more than raw throughput.

Originally published in Think Big Newsletter #35 on Amir Elion's Think Big Newsletter.

Subscribe to Think Big Newsletter