Agentic Engineering Setup (after 2,000+ hours)

Over the past 3 years, I've spent well over 2,000 hours coding with AI, and I've personally interviewed some of the most productive people in the space of Agentic Engineering.
Below is my full Agentic Engineering setup as it currently stands, Q3 of 2026.
Interface
This means the UI / CLI how you interact with agents. My main interface is bb.
It's open source, completely free, and it allows you to use any subscription, any agent, any model inside of a single GUI. Codex, Claude Code, Pi, Cursor CLI, OpenCode, Grok Build, Hermes, all within the same UI.
The problem with apps like Codex or Cursor is that they only allow their own models and their own subscription. The goal is getting the most tokens for the least amount of dollars.
All the features you like from the Codex or Cursor apps are inside of bb, and it's improving every single week (plus, it's fully open-source and 100% free to use).
Another thing I use a lot is cmux.
When you launch a new cmux workspace, you can divide the screen just like you would in tmux (that's why the similar name), launch different terminals in every pane, and there's a built-in browser.
Where cmux breaks is once you have a lot of agents and workspaces. The left sidebar really is not the right primitive. Fine for a couple of things going on, but for serious agentic engineering work at scale it's not the best.
I use Ghostty as my terminal because it's very fast and native.
Inside of Ghostty, you can run Herdr, which is basically tmux but for agents, a backend runtime for agents. Very minimal, very lightweight, lives in the terminal, and when an agent is finished, its state shows on the left: done, idle, blocked, running.
Tracking the states of AI agents is ESSENTIAL. My prediction is that in 3 to 6 months this "agent state tracking" becomes more and more important, because you won't be talking to a single agent. You'll be talking to a manager agent that manages a lot of worker agents.
The last interface I have to mention is Corral, something I developed myself.
Instead of randomly switching between agents whenever they finish (in Herdr, there's no real order), every agent has a priority. Just like tasks have different priority/importance levels. When a P1 agent finishes, he goes to the top. You should never respond to a P4 agent when a P1 agent has finished running. (yes... i do need to make corral open-source. haven't got to it yet.)
Models and Subscriptions
You want the most tokens for the least amount of dollars possible. This should be one the main goal of every Agentic Engineer (after getting shit done).
There are 4 main subscriptions right now, and yes... this could be completely different 2 months from now.
The best "bank for your buck" deal right now is OpenCode Go. It's only $10, and it gives you Kimi K3, Grok 4.6, GLM 5.3, DeepSeek V4 Pro, and many other models... But it doesn't have the best models like Fable 5 and GPT-5.6 Sol (plus, the usage limits are quite small)
So, if you have a bit more money, here's what you should do:
btw... the Cursor plan is very underrated. Cursor, aka Grok, is going to become a great subscription because of the SpaceX acquisition. SpaceXAI has loads of compute, so they're allowed to play the game of subsidizing. And -- I think -- the Cursor/Grok plan gives you separate limits for Cursor + Grok Bot, which is just incredible.
Grok Bot is rapidly becoming the new way people interact with agents, so having a Cursor subscription has never been more important (not sponsored lol, it's just true). Plus, the new model -- Grok 4.7 -- is right around the corner.
No matter what, do not pay API pricing. It is the worst deal out there. Just get the subscriptions.
Cloud Agents
It's quite obvious that cloud agents are the future. Cursor, Amp, Devin, Codex... all of these companies are betting everything on Cloud Agents.
The proof that cloud agents are the future is the graph below.
[image here]
This is Cursor's internal share of merged PRs from cloud agents: around 10-15% at the start of this year, approaching 60% now. And that's merged PRs, the stuff that actually gets used. Soon enough this will be 70%, then 80% and then 90%.
The issue with running all agents locally, on your machine, is that it's not scalable. You cannot run hundreds of agents at once. All it takes is a couple of agents deciding to run your entire test suite at the same time and your computer will begin to make weird noises (even my $7,000 macbook pro struggles).
Cloud Agents give you isolated environments, persistent sessions, durable access to internet and electricity. If you close your laptop, you lose your sessions. Lose internet for a couple of minutes, and the harnesses cannot self-recover.
The problem with the existing cloud agent solutions is the INSANE level of ecosystem lock-in. Setting up the environments and all your secrets is a lot of hours, and then you're locked in: your sessions are there, you're on their pricing, and you give them all your data. Even if they don't train models on it, there are so many other ways to use your data.
The solution is to have your own server. And thanks to AI, this takes like 10 minutes to set up (seriously). Just get a VPS, run Herdr on it, and SSH into it.
Herdr gives you persistent agent sessions, and SSH lets you connect from your phone, your laptop, anything. You can literally achieve the 80/20 of cloud agents for a couple of dollars, without the lock-in.
Build your own Cloud env
I use Hostinger for my VPSs, and a KVM2 plan is enough. Here's the main thing I want to get across... you don't need to be an expert in VPSs, DevOps, Linux, none of that.
Just talk to your agent in plain English!!!
In the upcoming video, I set the whole thing up live. A cmux workspace, a coding agent in the left pane (Cursor CLI running Grok 4.6), an empty terminal pane on the right.
My cmux skill lets the agent find that other pane and run commands in it. I SSH'd into the fresh VPS myself, then told the agent "learn everything about that server, and set up the dev environment. Herdr, Node.js, Python 3, Git".
It analyzed the VPS in seconds, installed everything, started Herdr, then installed Pi Agent, found an OpenRouter key on my macbook, and went through the full setup itself. A few short prompts later, I had Pi running GPT-5.6 Sol and a second session running Fable, both in the cloud, on my own VPS, with full root access.
If something happens to my computer or wifi... If my MacBook explodes, those agents keep running. The full walkthrough is in the video. Link to my YouTube here.
One more speed thing... I dictate with SuperWhisper. Most of you reading this probably type at 40 or 50 words per minute. That is very slow.
BUT! you can speak at 250+ WPM. A voice AI tool (like Superwhisper, Glaido, Whispr FLow) instantly makes you 3-4x faster at sending prompts. Use one. Don't be stupid.
Harness
First harness I need to mention is Pi Agent, the GOAT.
The most minimal harness out there: just 4 tools, always runs in YOLO mode, supports any model, any provider. Very elegant, super configurable, and that's why so many people build on top of Pi. It's open source, completely free, just go to pi.dev and get it. Non-negotiable. It's the first harness I put on the VPS.
Cursor CLI. Very underrated, because you can use all the models: Grok, GPT models, Anthropic models, Kimi. You can tag skills, and you can pre-send messages. Great harness overall.
The next category of Harnesses is what I like to call "self-improving" harnesses.
The two most popular ones are Hermes Agent and Prime Agent. This is for when you don't know what you're doing. If a task has a lot of uncertainty, a lot of figuring out to do, use a self-improving harness, because it creates skills and improves with you over time.
And finally, the classics, Claude Code and Codex. I have them as aliases. A lot of people type claude --dangerously-skip-permissions every single day. Extremely slow, extremely inefficient. I type cc and it launches Claude Code with permissions bypassed; cx launches Codex in YOLO mode.
YOU MUST create global aliases for the long commands you run often. That's one of the laws of agentic engineering: How can you get more done in the same amount of time?
Skills
My skills repo went viral last month. (see github.com/davidondrej/skills ) It's also completely free, open source, all that.
The SKILLs most relevant to Agentic Engineering are:
*(1) /total-review
IMPORTANT: anything you repeat often enough should become a preset.
If it's a single step, use text replacements. I have them as Raycast snippets. "Answer in short in plain English." "Make your previous answer simpler and shorter." "Stage all files, write a clear commit, push to GitHub."
If it's a multi-step workflow, turn it into a skill.
(2) /ask-then-build
(3) /deepapi
(4) Guardrails and push lock
Don't install all of my skills. Just grab the ones you need.
Worktrees
A worktree is basically a copy of your primary checkout into a separate folder and creates a new Git branch there, so agents can work in parallel, completely isolated.
On a small project, that's absolutely overkill. Stay on a single branch and work faster.
On medium-to-large projects where you're running 20-30+ agents at all times, there's no way to avoid it. Without worktrees, the agents will conflict, reverse each other's changes, and fight each other. Another good thing about BB: it has built-in worktrees. It remembers that on my big repos I always want a new worktree based off of origin/main.
Other Agentic Engineering Tips
Know when to use each model.
Know when to review.
Pre-sending.
Subagents are overused.
ADRs.
Tests.
Prod DB access.
Track your agentic productivity.
That's the setup as it currently stands. In a month, it's probably going to be different. This changes all the time.
by David Ondrej (spoken for YouTube, then re-written for article format)
