🤖 I gave myself a billion-token-a-day challenge


The billion-token challenge

A month ago I heard a CEO talking about how his company had reduced its AI usage from more than 42 billion tokens a month to around 30 billion.

I don’t remember who it was. I remember wondering how much I was using.

When I checked, it was already more than 100 million tokens a day.

My first thought was that I could be doing more. Actually, my agents could be doing more, and I could be doing less!

I’d been getting more disciplined about writing specs and automating the work around my agents. I wanted them to keep working without needing me to start the next task every time one finished.

So I gave myself a goal: one billion tokens a day.

I figured getting there would force me to automate the steps that still depended on me. I wanted the work to continue while I slept or spent time with my family, without having to check whether another session needed an answer.

There are easy ways to burn tokens. Give an agent an unclear task and let it go in circles. Or let agents become so pedantic that they spend more time debating minor details than finishing the work. That wasn’t what I wanted. I wanted to see how much useful work I could keep moving through the system.

I couldn’t sit there all day

I don’t have the time, patience or mental capacity (iykyk) to sit in front of an agent session all day.

Approve this. Explain that again. Correct the direction. Come back after a meeting and figure out where it stopped.

Even when the agent is doing good work, being available for every next step is exhausting. I have a limited amount of attention, and I don’t want to spend all of it supervising an agent.

Adding more agents doesn’t solve that by itself. It gives you more sessions waiting for an answer.

I needed to put more of the direction into the work before it started. That meant writing specs that explained what I wanted, what was out of scope, and how to tell whether the result was correct.

Then I needed the handoffs. An implementation goes to a separate reviewer. Findings go back for revision. A maintainer follows the pull request through checks and unresolved comments.

I still have to decide what matters and make decisions the agents can’t make for me. But I don’t want every routine step to require my attention.

Could the work keep moving when I wasn’t sitting there? Each agent needed to know what it was allowed to do, which checks had to pass, and when to stop. I also needed a record of what happened so I could investigate when something went wrong.

Getting enough model access

A billion tokens a day turned out to be hard to reach with the mix of services I was using.

My setup includes multiple $200-a-month subscriptions across OpenAI and Anthropic, along with free models on OpenRouter. I also run a Qwen model locally on a MacBook Pro with 128GB of RAM and a DeepSeek model on an old gaming system.

When a capable stealth model becomes available on OpenRouter, I use the opportunity while it lasts. Ox Alpha turned out to be Z.ai’s GLM-5.3-Flash. Union Alpha was revealed as Unbiased Pareto. More recently, Space Bunny Alpha accounted for a large share of usage.

When there’s no suitable stealth model available, I use other capable free models, such as Nemotron.

But free access comes with questions I need to answer before sending work to a provider. Most of those questions apply to paid models, too.

What am I willing to send it?

With a stealth model, I may not know who’s behind it. Guessing its identity from how it responds doesn’t tell me where my data goes or who can retain it.

That’s a problem if I’m about to give it access to private source code.

​Space Bunny Alpha’s listing says prompts and completions may be retained, but aren’t used for training. Those are different promises. Ox Alpha’s listing says the provider retained them, too.

A free model can be useful for public-code research or a task that contains nothing sensitive. That doesn’t mean it belongs on every task in the queue.

And this isn’t just a concern with an anonymous model or a provider based in China. An American provider doesn’t get unconditional access to my credentials or intellectual property either. I still need to understand the terms and decide what I’m willing to send each model.

Before the harness chooses a model, it needs to account for what it’s about to send, not just what the model costs.

Otherwise, it’s very easy to save money on inference by sending something somewhere it shouldn’t have gone.

Building the tools I needed

I’ve been building Alyria, an AI agent platform with a cybersecurity focus. Beacon handles agent governance. Statio is the orchestration work.

Trying to keep my own agents working gave me a reason to use what I was building every day. It also made the requirements specific.

An agent might need to push a commit. That doesn’t mean the model needs to see a GitHub token. I want the execution layer to supply credentials only to the authorized operation, without putting them in the model’s context or exposing them through tool output and logs.

I also need to choose which providers can handle which work. A model whose provider retains prompts might be acceptable for a task involving already-public code. The same model shouldn’t automatically receive a private repository, internal documentation, or the conversation history from another task.

And public code still needs review. Keeping secrets out of a prompt doesn’t prevent a model from generating vulnerable code, including code that could expose data when it runs.

I want a separate model reviewing the changes, automated checks, and explicit permission boundaries around actions such as pushing or deploying. Another model’s approval isn’t a security guarantee, but it gives me another opportunity to catch a problem before the code runs.

These were the controls I needed before I could comfortably leave more work to the agents.

Have I met the goal?

Yes. I’ve passed the daily token goal for 35 days in a row. Some days, I’ve more than doubled it.

For the 24 hours from October 3 to October 4, 2026, ending at 1:25 p.m. Eastern, my agent harness recorded 2,219,932,407 tokens.

Space Bunny Alpha accounted for 1.23 billion tokens at a recorded cost of zero. Without it, the rest of the models accounted for about 985 million tokens. Having a capable free model available made a substantial difference that day.

Now that I have officially hit twice my goal, what's next? I want to keep improving how much useful work gets done without needing more of my attention.

I’m Brian McManus, and I’m building Alyria to help define, run, and govern AI agents, including what they can access and which actions they’re allowed to take.

If you’re working through the same challenges or thinking along similar lines, drop me a note. I’d like to hear what you’re building. You can also subscribe to my newsletter at brianmcmanus.com.

Brian McManus

Build Without Limits is where I share my journey—an unfiltered look at how I challenge limitations, break conventions, and explore the strategy behind great software products, with occasional insights on AI. Subscribe for real, unfiltered lessons on entrepreneurship, problem-solving, and innovation.

Read more from Brian McManus

Brian McManus August 16th, 2026 No Plan Survives First Contact What preparing an autonomous agent for HalCTF at DEF CON 34 taught me about model routing, harness design, and adaptation Seventeen years after my last DEF CON, I knew almost immediately where I would spend every waking hour: the AI Village, competing in HalCTF. In a traditional capture-the-flag competition, or CTF, you hack a vulnerable system and score points by extracting flags: secret strings that prove you got in. HalCTF, the...

I just got back from W3C TPAC 2025 in Kobe, Japan. I participate in the Web of Things working group. This year’s Plugfest (European word for hackathon) brought together dozens of companies building solutions for the open web. I left consumed by one thought: most builders are being restricted by two kinds of constraints. The ones they don’t know they have. And the ones they refuse to set. The Invisible Prison The first kind, unconscious constraints, is what my newsletter is about: the...