A little analogy that has been translating quite well for me in an age of agent-driven development comes from my days of working in corporate.
We’ll be nerding out a bit today. Are you ready?
Some context
For those who’re new to this Substack (we have a few new people in the house, hi!), I used to work as a Product Manager at Atlassian till this year. My job was to work with a team of developers, designers, GTM, FP&A people1 and ship some platform features, say the ability to create a partial refund on our billing system which is used by nearly all Atlassian products today.
Now, to get these features out of the door I had to be in constant engagement with the engineering team. The team was quite large though, at least a team of 10 cross-functional devs who managed different things at different surfaces. Which is why my discussions were 80-90% limited to the most senior engineer (called the tech lead (TL)) in the team or their reporting manager (called the engineering manager (EM)).
The TL/EM were supposed to be the directly responsible individual (DRI) for a feature’s implementation. Something broke? You went right to the TL, who will coordinate with the right engineer and we’d together sit with them to get things sorted. This often looked like a few hours of meetings and code changes at its best. At its worst, it turned into days of meetings, escalations, an infinite matrix of Slack conversations and de-prioritization because there were other important features each of us had to deliver too.
It’s not that the engineers or I were getting slow on purpose, it was simply because of our very human communication constraints in a fully-remote work setting. It’s also not that the TL can’t go into the system and fix it themselves to be quick (these days it’s called high agency™️), it’s just that their attention was spread too thin across various dependencies, and they lacked context of the deeper details to be of too much immediate help.
The EM will have an even lesser context per-project since they’re coordinating engineers across projects. The VP of Engineering: even lesser context into project-level nitty-gritties since they were managing workstreams. Each of them traditionally get paid more for breadth, more experience and more responsibility.
That’s hierarchies in large tech companies for you, in a nutshell. In any sort of large company, more broadly. Be in the world of finance, bureaucracy or law – you will relate to it.

In my experience of working with GPT-6 Astra this week, I experienced a similar hierarchy mushrooming with my agents.
For those who still don’t know (as I’ve purposefully been not too loud about it): I’ve been working solo on this little project called saras.works. Think of it as Muse but specifically meant for your GMAT prep. Not here to advertise about it here, so moving on...
There are A LOT of moving parts to Saras. Most of things were things I didn’t know about at the slightest some six months ago. But now, I’m peering at the frontend, backend, QA, design, growth, model behaviour, memory considerations, cloud deployments and so on. Not on the same day typically, but at least one rotation across everything once a week.
My eyes are watery, my screen time’s almost the same as my time awake, my back’s hurting and my attention has officially started to fracture. Which is why I’ve had my first hire earlier this week. But my persistence to keep on operating as a solo founder on a product which obviously required multiple engineers has come for a few reasons.
Firstly, I am enamoured by the very fact that I can keep on doing all this as a solo founder, even though I didn’t think it was even physically possible at least till the June of this year. Secondly, going deep into these systems myself allows me to see what is the kind of org structures will come to be in the post-AGI era, letting us decide the due course of action for our company. Thirdly, keeping humans (which also means myself) out of the loop has a few benefits of its own.
When Fable 5 was released for the first time by Anthropic (back in June), I couldn’t find the right occasion to use it. I felt I didn’t have anything big enough to throw at it at that point in time. I could easily use an Opus 5 (and dare I say it, Sonnet 4.5) on my problems and get the same tasks done for much cheaper. For me, doing the task and doing it cheaply mattered more as compared to using the best model to do it. My job is to keep my customer happy, and my customer doesn’t really care what did I use to build it as long as it does the job for them.
With each improving model, the marginal utility started feeling smaller. New models weren’t really impressing me anymore.
I couldn’t really process the hype behind Fable 5. I won’t lie when I say that when Anthropic revoked early access to that model, I felt lowkey guilty of not being able to make the best use of it. But still, for the life of me why would I waste tokens as a pre-revenue founder?
But I realized my perspective of looking at the trajectory of improvement was very uni-dimensional. My mental model didn’t let me grasp in exactly what metric these models were improving and how can I make the best out of their rapid improvement until very recently.
We need to come back to the corporate engineering org structure diagram that we were discussing earlier.
Long story short: We need to start thinking of frontier models as a hiring a leader than hiring a junior/intern. Allow them to help you out with the coordination of overhead of having to deal with multiple agent sessions, which can run on cheaper, smaller models.
The direction of AI progress is not simply on a single dimension – faster, cheaper, accurate. Now that these problems are reasonably solved, we’re improving these models on a different dimension altogether – their ability to run swarms of agents and attacking at different aspects of the same broad problem in sub-steps.
We saw a first glimpse into this approach of breaking down problems when GPT-4o came out and we were still figuring out how to make a model count the number of r’s in the word strawberry.

Subagents are an extension of this idea. Fresh context shared across multiple agents, each of which are responsible for a single deliverable.
You’ll greatly inhibit yourself if you’re a part of the single-shot club. The new models can surely take you the extra mile in building yourself a clone of Doom, but that’s not they’re the best at. The new models may not feel cheaper or faster. But they feel highly accurate and agentic, which is another way to say they can hold a lot of information in and act on it across other agent sessions.
I want you to look at all the words I’ve highlighted in the paragraph above.
Go on, I’m waiting for you.
All of these problems that I have described above with people (deficit of attention, lack of context, cost) also happen to be very agent-shaped problems. Like humans, agents also run out of working memory. They also start hallucinating when 50% full. They start costing more with every query that they make. And if you’ve ever used them, you’d know they cost a bomb and run your weekly usage dry in just a few prompts.
The larger your knowledge base is, the more amplified each of these problems become. The worst you can do is throw the entire directory of 100s of files right at the agent, give it the vaguest prompt imaginable and pray for the best. The output often is like a one-pot meal with all the veggies, spices and oils without any proportion.
Distasteful.
Eww.
Something that the frontier model labs may hate me for saying, but it’s true. You don’t need a frontier model to do your basic tasks, and that’s okay! As much as the labs would promote their newest state-of-the-art model for running “more accurate Excel analysis” from your spreadsheet, you’d be better served by a cheaper, open-weights model. Sometimes not even that – a free lightweight ML model that runs natively on your machine would be more than sufficient to improve the quality of your voice recordings or upscale your thumbnail for your next podcast project.
Simply put, the more specific the problem, the smaller unit of intelligence you need.
Which also means that you don’t have to use GPT-6 Astra to go through each row in your Excel file to get some tedious work done. You just need a smaller model with a better harness.
These large models may create a temporary “harness-like setting” for you, to run lightweight agent swarms, and clean after themselves when they’re done. That’s their capability.
Drawn in human terms, they’re like fractional leaders that you appoint to an org to restructure the company to get some specific outcomes out, resolve conflicts and see off when the time is right. All this, but at a much smaller scale for specific tasks spanning 1-2 days at a time.

Think about it – you wouldn’t hire a PhD and make her do basic data entry work, would you? That’s the neither the best use of her talents nor it is the best return on your limited budget.
But you sure need an “expert” for doing the kind of work that an expert could. The only change in mentality is to allow the expert to “hire” and build a team around her. These “hires” are typically in the form of cheaper, faster models.
This approach has the following benefits:
more work done in parallel because each agent works in its own lane
lesser coordination overhead between different agents for you
lesser chances of error (surprising, but true)
more rounds of validation before anything goes out => higher quality of output
much better results for much cheaper (even cheaper than running cheap models and committing cheap errors)
So let the intelligent frontier models be the leaders. Hire them for coordination and validation, not the grunt work. Enjoy the benefits of abundant intelligence and chase after your more ambitious goals.
These goals can be big or small – getting your startup off the ground, negotiating a higher pay at your job or just earning the liberty to walk away from your desk and making yourself a glass of Masala Dew in this hot/humid weather.
I tried simplifying things to the best I could, but I also understand some of it may sound like an alien language to you. However, if you found this post helpful and you’d like me to cover something else for the next week, do let me know.
A little pitch
Based on all I’ve learned in
my few years of working in big-tech and shipping features worth millions of $
the last few months of keeping up with developments on the frontier (it’s more than a full-time job!)
building outcome-oriented agents for kind folks outside of the frontier tech bubble
So I’m opening up a bit of time for consulting, talks, and other interesting collaborations. Whether that’s helping your team adopt agent-driven workflows or just thinking out loud together.
There’s only one catch – we like to keep it fun!
If you’d like to have a chat, please don’t hesitate to reach out at samyak@sanganak.works
Have fun building new things. But above all, stay hydrated.
Cheers,
Samyak
if you don’t know GTM and FP&A, don’t even bother. If you know, well you’re in my world already.





