On building · AI-assisted engineering
Vibe DevOps: the boring parts that make it work
Invest your discipline at the boundaries — inputs, outputs, contracts, guardrails — and let the model be loose and creative inside them.
Whatever vibe coding turns out to be — fancy autocomplete or something genuinely new — people tell me I get more out of it than most. I don’t think I’m doing anything magical. I just think the leverage is somewhere different than where most people are looking for it.
Everyone focuses on the prompt, the instruction, the conversation with the model. That part matters, but it is not the bottleneck. The bottleneck is everything around the code: how you know it is right, how you know it still works tomorrow, how you ship it, how you secure it, and how you keep it standing when something fails at 2 a.m.
I’ve started calling this Vibe DevOps. The idea is straightforward. Invest your human discipline at the boundaries of a solution — the inputs, the outputs, the contracts, the guardrails — and then let the model be loose and creative on the inside. The tighter your boundaries are, the more freedom you can afford in between them.

01The eval is the deliverable
The eval is the single most important human input in the entire process — more important than the prompt, more important than any architecture sketch. Concrete examples of “done” are what close the gap between what you imagine and what the model produces. If you spend your time on one thing, spend it here.
The more examples you provide of your definition of done, the more likely you are to get there. When you miss something, add it to the eval and run again. This is not a side artifact or a nice-to-have. It is a first-class deliverable — arguably the primary one.
I suspect there is a theorem somewhere that a sufficiently detailed eval maps to a sufficiently correct implementation. But even the intermediate states are helpful, because each pass tightens the loop between what you asked for and what you got.
02Unit tests are for future you
The point of unit tests is not to verify the code you just wrote. It is to protect you from the code you are about to write. When an AI assistant rewrites a module, it does so confidently, plausibly, and possibly wrong. High test coverage is the safety net that lets you move fast without moving recklessly.
This applies to smoke tests, integration tests, and edge-case tests too. Coverage is not a vanity metric in this world. It is your insurance policy against a collaborator that does not remember what it did five minutes ago.
03Validate at the seams
If the model is writing your data pipeline, it is going to be creative about how it structures what flows through it. That is fine — as long as you are strict about what goes in and what comes out.
Pydantic is your best friend here. Define your data models explicitly. Validate inputs on the way in and outputs on the way out. If the model hallucinates a field name or quietly changes a type, Pydantic catches it at the boundary instead of letting it propagate downstream into something much harder to debug.
This is the “discipline at the boundaries” principle expressed directly in code. You are not micromanaging the internals. You are enforcing contracts at the seams.
The tighter your boundaries, the more freedom you can afford in between.
04Docker, all the way up
Work locally, push to staging, push to prod. Same container, same environment, same behavior. This is valuable on its own, but it becomes especially important when half your stack was written by a model that does not remember what it did yesterday.
Version your images. Tag every build so that when something breaks in production, you can roll back to a known-good state in seconds instead of minutes. Using latest as your deployment strategy is not a deployment strategy.
Add health checks to every service. Docker’s HEALTHCHECK instruction is trivial to configure and gives your orchestrator the information it needs to restart containers that have gone sideways. The model will not think to add these on its own, so you need to make sure they are part of your process.
05Build for resilience
AI-generated code tends to solve for the happy path. It works when everything goes right, but the question you need to ask is what happens when things go wrong.
Think about failure modes early. What happens when a dependency is down, when an API returns a 500, or when your queue backs up? The model will write retry logic if you ask for it, but it will never spontaneously ask “what if this fails?”
That question is your job. Timeouts, circuit breakers, graceful degradation, and structured error handling are what keep a system running in the real world — and they are exactly the things that get skipped when code is being written by something that optimizes for “does it work right now?”
06Security is not a phase
Security needs to be part of the process from the beginning, not something you bolt on at the end. How much access does this thing need? Should it be screwed down or left open? What are the expectations for where this code is going — an internal tool, a staging environment, or something customer-facing?
You need to answer those questions early, before the model has already scaffolded an architecture with wide-open defaults. Get in the habit of running /review and /security-review in Claude Code. It takes a minute and it is worth it every time.
The model will not think adversarially about your code unless you make it do so.
07Git discipline
The model generates large diffs. It rewrites whole files when you asked it to change one function. If you are not careful about version control, you will lose track of what changed and why — and three days later you will not be able to figure out which commit introduced a bug.
Small, frequent commits with clear messages. One branch per feature. Review the diff before you commit, not after. The model does not care about git hygiene and will never ask you to commit before moving on, so you have to build that habit yourself.
This is also your undo button. When the model goes sideways and you need to back out of a bad direction, clean git history is what lets you do that in thirty seconds instead of thirty minutes.
08Automate the pipeline
If you are manually running tests, manually building Docker images, and manually deploying, every one of those manual steps is a step you will skip when you are tired or in a hurry. And you will be tired and in a hurry.
GitHub Actions or any CI/CD tool can turn the whole chain into something that runs without you touching it. Push triggers tests, tests trigger build, build triggers deploy. You set it up once and then it enforces your standards whether you are paying attention or not.
This is where all the other practices come together. Your evals run automatically on every push. Your test suite catches regressions before they reach staging. Your Docker image gets built, tagged, and deployed with health checks already configured. The pipeline is the thing that makes all the other practices sustainable, because it takes discipline out of your hands and puts it into the infrastructure.
None of this is about distrusting the model. It is about building the kind of environment where you can trust it more — where its creativity has room to work because the guardrails around it are solid. The discipline is not in the code itself. It is in knowing what good looks like and being able to prove that you got there.