Exponential View

Exponential View

🔮 Seven lessons for managing AI agents

Plus, an updated stack of 50+ AI tools we use at Exponential View

Aug 05, 2026
∙ Paid

In April 2025, we shared our seven lessons for building with AI. Many still hold. But agents have changed how we work, so the lessons deserve an update.

Agents can now work on harder tasks for longer. They plan, use tools, work without human oversight, and act on our behalf. In May, roughly a quarter of Codex users were making at least one request per month for work that would take a human eight work hours to complete. This is up from 2% in December 2025.

Our role as managers of agents is evolving with the models. There is no playbook, so experimentation is still the best way to learn how to get good at it.

Our team recently sat down to review what we’ve learned from working with AI agents over the past six months — today’s seven lessons are distilled from this team meeting.

We’ve also updated our internal stack of 60+ tools – everything we’re actively testing, using or intend to use. Become a member to get access to the full stack.

Upgrade for access


1. Write the finish line before the goal

AI agents are sometimes too eager to declare their work complete, even when it’s far from done. It doesn’t mean that AI is “lying”; it may have misinterpreted your goals. And if you never specified the end goal, it pretty much just guessed it.

To set yourself up for a good autonomous run, before you do anything else, write a finish line to answer one question: “How will I know this is done?”.

Azeem is a big proponent of handwriting to help him think, and this would be the right time to use your pen and paper to think through what you expect to see at the end of the run.

Agents become much more useful when “good” or “finished” is something they can test – in our experience, evaluative finish lines will get you farther than descriptive ones. A simple example, instead of ordering your agent to “make this Rubik’s Cube look more organized,” instruct it to “solve the cube; every face must be one color.” As AI gets its most intensive training in coding, we try to recreate similar environments in our tasks.

Let’s say we want to task ChatGPT with building a small Python module to process work logs. It needs six functions, each checked against six tests, a total of 36 tests to show it’s done the work.

First, how not to do it:

Build a Python module for processing work-log data. Implement these six functions […] Reply with the complete module, and say `STATUS: COMPLETE’ if you believe it’s ready.

A better finish line would be explicit and testable:

FINISH LINE – Do not claim completion unless the full 36-case test suite passes under Python 3.14. All six functions must work, imports must succeed, inputs must remain unmodified, and only the standard library may be used.

You can use the same rule for non-engineering tasks. It may be trickier, but not impossible. Show the agent what a completed deliverable needs to look like, or give it a pre-filled template, as recommended by Anthropic’s Applied AI team. Your instruction for such a task may look like this:

a 1,200-word memo for a board deciding whether to approve an AI-infrastructure partnership; decision and three reasons on page one; every material number linked to a dated primary source; facts, estimates and assumptions separated; base, upside and downside cases; the strongest contrary evidence represented; stop and escalate if two material sources cannot be reconciled.

Some tasks won’t be right for agents. We were recently exploring a project to build a network of beliefs and relationships, but not really knowing what a useful final output would be. This was not a good candidate for a long autonomous run – so we first spent time clarifying the goals before we assigned an agent a task.

2. Spend intelligence where it changes the outcome

A year ago, before prompting AI, we’d have asked: “What’s the best model to use for this query?” Today we’re more likely to ask, where in this workflow does additional intelligence change the outcome?

You don’t need the most capable model like Fable 5 performing every step in your task. It will be slow and expensive. We’d use cheaper models to do the grunt work. Our OpenClaw agents run on DeepSeek V4 Flash most of the time.

For some tasks, however, you’ll want to start off with a strong model right away. Let’s say we’re investigating Europe’s compute shortage outlook. Before we dispatch agents to collect evidence, we’d deploy a stronger model to set research parameters first, define what “shortage” means, decide the forecasting horizon, and set out rules for how conflicts in research will be resolved. Once we’re happy with the framing, cheaper models can go off and do the work.

Effort is one of the levers you’ll want to use to adjust intelligence per task. In one benchmark, GPT‑5.6 Sol improved from 49 at low effort to 59 at maximum on Artificial Analysis’s Intelligence Index. Yet the final stretch, jumping from xhigh to max, doubled output tokens for a one-point gain. More effort is not always better value.

The rule of thumb from Anthropic’s recent lecture, which our team attended, is to prefer a larger model at low effort over a smaller model at maximum effort. More model before more effort.

3. Leverage over token count

Azeem hit his first 100 million tokens-a-day mark in February. OpenClaw completely changed the way he worked. He estimated that one overnight run was equivalent to 48 hours of his work time.

The token count is one way to measure how we use AI, but it doesn’t measure the quality of work. Tokens are a bit like electricity in a factory, measuring what goes in but not what comes off the production line. In our State of the AI Economy report, we proposed a quality-adjusted output token as a better unit of value:

Until there’s a better unit of value for intelligence, you can use approximations to understand how good of a colleague your agent is. We recommend a light weekly audit of the substantial tasks AI attempted, which outputs you ended up using, your model and infrastructure costs, the time you spent briefing and reviewing the work, any corrections or reruns – and the estimated human-equivalent hours.

Azeem’s first audit back in the spring showed that over the course of one week, his OpenClaw agent performed 62 substantial tasks and incurred costs of about $800. He estimated that commissioning the same work from humans would’ve cost him around $19,000 and 48 hours of his time. It’s an estimate, sure, not an accounting-grade ROI. But even a light audit will show you where your agents have most leverage.

4. Don’t argue, restart

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 EPIIPLUS1 Ltd · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture