All issues
Issue 2425 Sep 2026By Cara Davies

AI you can stand behind.

AI can now produce more than any of us can check. The problem is now making sure that we can stand by the output our agent is creating.

Hi folks,

This week I have been thinking about the difference between AI that makes more work and AI you can stand behind.

What we're seeing.

AI you can stand behind.

Shane Parrish on X: "Lazy work used to mean too little output. Now, with AI, it often means too much and more work for everyone else."
Shane Parrish on X.

Shopify's Tobi Lütke has a name for it: the slop grenade, a term he credits to Harry Brundage. Shane Parrish wrote it up this week: you let AI produce the work and pass it on without adding any value, including checking it. Someone else has to wade through it and clean up the mess.

So what is the right level of AI? For us, it is AI that creates things we can stand behind. Standing behind something means we sign our name to it and feel proud of it.

Something I have noticed in my own work: when I cram a lot in, the quality goes down. I cannot stay on top of all of it and truly sign it off. If I cannot stand behind the work, what value have I created?

Working with agents is like being a manager. When a direct report gets something wrong, it is often the manager who is responsible, because the manager is accountable for the quality of the team's work. The US Treasury Secretary said as much on Monday about OpenAI's agents breaking into Hugging Face: "that is the responsibility of the OpenAI management, not a bunch of agents."

AI can create so much, so fast, that we humans cannot always keep up (it is exhausting to try). So we need systems that keep the team's work quality where it should be, and the team now includes AI agents.

What we are building at Levercon is AI you can trust and stand behind:

  • Human approvals at the key points. We take the way the work is done today and make sure the approval points and check-offs are in the right place, so people can truly stand behind the AI's output.
  • A clear trace for every agent. For each agent we build, you can see how it got to its answer, so as the agent's manager you can sign off on the work.
  • Feedback that sticks. We make it easy to give feedback and have the agent improve from it.
  • Explainable by design. A large part of our thinking goes into making the AI explainable, and that thinking is still evolving.

This is a constant work in progress in how we build our products!

In the mix.

  • Claude Opus 5.5 is out (Anthropic)
    • Released on Tuesday. Anthropic says it performs at the level of its top model, Fable 5.1, on most work and costs 40% less to run than Opus 5.
    • It is faster at building merger models in Excel. Anthropic also tested whether it would invent numbers when writing up a company's quarterly results: 16 of 18 reports were clean, 2 had something made up.
    • Why it matters: that is a strong analyst who is wrong once in every nine reports. You would not send their work to an investment committee without reading it, and the same goes for the agent.
    • My vibe check: it is much faster, so fast, which makes it really good. It also talks to itself much less and just does the job, which is nice. It does mean it can sometimes get a bit too goal-seeking in how it tries to reach the goal you have set it.
  • Jev, a new kind of AI model that is not an LLM (TypeSafe AI)
    • Jev is not a large language model (LLM). Models like Claude and ChatGPT write text, one word at a time. Jev does not write text at all.
    • You define the possible answers in advance, and it returns one of them with a probability and a confidence score. It works more like a classifier.
    • Because it can only pick from answers you set, TypeSafe says it cannot hallucinate.
    • It is fast and cheap: TypeSafe claims 70 to 500 milliseconds a response, and it does not charge for output.
    • Why it matters: one of TypeSafe's headline uses is using Jev to check other AI's output, scoring and verifying an LLM's reasoning and outputs. We believe this could be a fast, cheap second pair of eyes on an agent's work, at exactly the check points we care about. We will be experimenting with Jev over the next few weeks.
Written by
Cara Davies
Cara Davies
Director | Product & Engineering

Levercon builds custom, self-improving AI agents and systems for investment firms, with every output traced, verified and signed off.

Receive the Weekly Brief every Friday in your inbox.

Get Started

Book a strategy call.
See where AI fits in your firm.

A short conversation about where AI could help your firm, what to build first, and what it would take.