The agents got better, and my job changed
It's been just over a year since my Tips for AI-Assisted Coding blog post, and it's time for an update because things have changed even more than I could have predicted. Most importantly, the models have gotten much better. They are capable of handling much longer and more complicated tasks now. This has changed how I use them. I no longer treat them as assistants or grunt workers. Instead, I think of them as a full engineering team at my disposal, and I'm learning to trust them with large projects. My job increasingly consists of deciding what I'm trying to achieve, providing the agent with the context it needs, and making sure I understand the result.
I've stopped using AI merely to write code. Instead, I use AI agents to brainstorm, learn new things, plan complicated systems, and understand existing ones. At times, I use them as if they were my interns, but at other times it's better to think of them as expert consultants. They are intimately familiar with all kinds of programming language quirks. They can teach me more about AWS than the average professional AWS consultant can. And if there's something they don't know, they can read the documentation at superhuman speed. Even when I'm confident in my high-level design, I know I will almost certainly benefit from asking the agent for an honest critique. There's a good chance it will poke some holes in my plans or bring up edge cases that didn't occur to me. In fact, writing code is no longer the most valuable thing AI does for me. Yes, it happens to be really good at it and a lot faster than I am. However, even if I had to do all the grunt work myself, I would benefit massively from the agents' ability to understand, explain, and design complex systems.
Most of my work now happens before implementation
My workflow has changed significantly since my blog post a year ago. I now spend most of my time brainstorming and reviewing specifications. Last year, the focus was on implementation planning and techniques for coercing the agent into writing decent, functional code. I even shared some tips for getting the agent unstuck when it inevitably painted itself into a corner. One year later, the agents have gotten so good at implementation that I no longer ask for a separate implementation plan. Even the post-implementation bug-fixing step has almost disappeared. All of this depends on being thorough during the brainstorming and specification stages.
Brainstorming is where I use the agent to figure out what I want. I ask it to interview me about the idea, helping me identify important decisions and discover edge cases I hadn't considered. Often, we brainstorm individual features, but sometimes we tackle whole apps or complicated migrations. If the scope is large, I tend to start with a session where we break it down into smaller parts to explore in later brainstorming sessions. I consider the brainstorming session successful when I can end it with a request for a specification or a list of new items to brainstorm later.
Reviewing the specification is my final check that the agent and I agree on what we're building. It needs to be detailed enough that I can spot any misunderstandings before implementation begins. One of my most important tasks now is to open the spec file in my Markdown reader, make a cup of coffee, sit down, and carefully go over the spec to make sure that my agent and I have reached a common understanding. It takes a little bit of experience to know exactly what needs to be specified and where you can just trust the agent to do the right thing. Also, the models are always changing, so I have to keep adjusting. Quite often, the spec doesn't need any revisions because all the decisions were already made during brainstorming. Still, I find problems often enough that I have learned never to skip this step.
The agent needs the context that lives in my head
As I delegate more of my work to the agent, I need to give it context that previously only lived in my head. Does the app have paying users? Is it internal to your company? What kind of infrastructure does it run on? Have there been earlier versions of this app, and what did you learn from them? These are all things that influence how you would architect a system. For example, if your app doesn't have users yet and you fail to inform the agent about this, it will probably plan an over-engineered zero-downtime deployment for your new feature. The frontier models quite reasonably assume that your app is critically important and cannot tolerate any downtime unless you tell them otherwise. They're playing it safe, which is good, but it has implications that you should take into account.
Beyond the current state of the project, agents need to understand the decisions that led to it. The codebase alone doesn't explain why something that looks like a bug is intentional or which approaches I've already tried and abandoned. This is knowledge that I would normally accumulate in my head over time or, if I'm lucky, record in a text document or my team's Slack channel. I now ask the agent to record those decisions and the reasoning behind them so that later on we can build upon what we've already learned. Writing decisions down would obviously have been a great habit even before the age of AI, but I always lacked the discipline. Agents help me both maintain a record and use it effectively.
To keep track of accumulated knowledge, I use a project management system that the agent can access. For small projects, I use a flat list of Markdown files, but for my largest projects, I have implemented custom project management systems. I organize the system around tasks, which have two important attributes: a status (in progress, done, etc.) and a list of progress notes where decisions are recorded. For larger projects, I also include commit hashes and links to specification files so that the agent can reconstruct the full history of a feature from its task file. Lastly, I write a skill explaining how to use the system and what to record, and I use it consistently in every session. In my experience, agents maintain this record reliably as long as I explain the conventions and ask them to follow them.
I also use the agent to understand and check its work
To lose understanding is to lose control. Agents can generate code faster than I can read it, but I still need to understand enough to make decisions and direct their work. However, reading every line of code sets a hard limit on how much I can achieve. I find it useful to think of AI agents as a team of engineers working for me. I can ask them to generate reports, write automated tests, or conduct code reviews, just as I would if I were leading a team of human engineers. When the situation calls for it, I can still open my IDE and do it the old-school way. I also still glance over diffs before merging changes.
Because code is now cheap, every new investigation can have its own custom dashboard. I use this new capability to understand not only data but also my codebase. You can ask the agent to visualize the dependencies between different modules or generate a map of your cloud infrastructure. With a large enough budget, all this is possible without AI, but now you can do it on demand.
I also use agents to investigate problems by analyzing application logs. I simply tell the agent what behavior I'm trying to understand. If the logs are missing information, I ask it to add logging and reproduce the problem. It will then trace events, piece together the evidence, and explain what happened. I can even ask it to visualize the results. I find that I'm now using the logs to diagnose issues that would previously have required too much manual work.
AI code reviews have become routine for me, just as I anticipated a year ago. They are cheap enough that I ask the agent to review every change before I merge it. When I review code myself, it's mostly to improve my understanding. I increasingly rely on agents to find bugs because they do it better than I do. Asking them to review code I've written myself gives me a tinge of anxiety because I know they will almost certainly discover a long list of issues, some of which may be critical. I consider myself a decent programmer, but it's time to acknowledge that AI can write code with fewer bugs than I can, provided that I let it review its work and fix any problems it finds.
Once a feature is implemented, I ask the agent to smoke-test it through the user interface, even if it has already written automated tests. This often catches problems that the tests missed. This has become practical over the past year as agents have gotten much better at operating GUIs. For example, a dropdown option may not be visible even if it's programmatically accessible, or a CSS bug may prevent you from scrolling all the way down the page.
What this looks like day to day
Because my agents now work on bigger tasks, I can run many of them in parallel without being overwhelmed. While my agents are working in the background, I focus on brainstorming sessions and reviewing specs, which I can only do one at a time. In practice, I run tmux on my laptop and create a new tab for each agent. I try to name the tabs based on each agent's current task so that I can easily identify them in the tmux tree view.
I use worktrees when I'm working on something I want to keep on a separate feature branch. That way, I can have multiple agents working without them stepping on each other's toes. Quick tweaks and bug fixes go directly on master, and I usually make sure that agents working on master handle separate parts of the codebase to avoid conflicts. That means there's still quite a bit of manual session wrangling involved. I expect this to change in the future, and there are, of course, a variety of tools for multi-agent work. I just haven't yet found anything I like.
Conclusion
A year ago, I spent more time figuring out how to get the agent to implement things. Now I spend more time figuring out what it needs to know and what I need to understand to maximize the amount of work it can take on. Understanding is where the new bottleneck lies. Right now, the only way I can get more work done is to sacrifice understanding and just trust the agent to do the right thing. However, that means relinquishing control to AI, which feels wrong, especially since I'm still being paid to do a job. That being said, things keep moving fast, so who knows where we'll end up a year from now?