00How to use this handbook
This is a book, not an article. Nobody is expected to read it in one sitting.
This handbook describes how the best WooCommerce developers work, and it is the standard we hold ourselves to at the agency. It is long on purpose. Every chapter is meant to be read, tried on real work for a few days, and then returned to.
What to expect
- Part I is about how top-tier developers think and work: solving problems, writing code, reading unfamiliar codebases, releasing safely, recovering from disasters, estimating, and communicating.
- Part II is about how we train for it in a world where AI writes most of the code, including the program, the review rubric, and the drills.
- Each chapter ends with a drill. The drills are not optional reading. They are the part that changes how you work.
How the reader works
- Your place is saved in this browser. When you come back, you will be offered the chapter you were reading.
- Use Mark as read at the end of a chapter to track progress in the sidebar. It only tracks; it does not judge.
- Use the arrow keys or the buttons at the bottom to move between chapters. Press / to search chapter titles.
- Switch to Scroll view in the top bar if you prefer to read everything as one long page.
- The Aa button opens reading settings: text size (also Ctrl++ / −), five themes including Paper and Dark, line spacing, and reading width. Everything is remembered on this device.
A suggested pace
One chapter per working day. Read it in the morning, look for it in your work that day, and note one thing you would do differently. That is 22 working days for the whole book, and it will do more for you than reading it all in one weekend.
Read chapter 1, the seven mindset shifts, first. Everything else in the book is one of those shifts applied to daily work. If you only ever read one chapter, read that one.
The traits, and what they look like in practice
It starts with the mindset shifts that everything else grows from. Each trait after that is covered in three ways: the idea behind it, what it looks like on a real WooCommerce production site, and a drill you can run with the team.
01The seven mindset shifts
Every trait in this handbook grows out of one of these shifts. Many developers never make them, and it shows in their work for years.
A junior and a top-tier developer can know the same hooks and the same APIs. What separates them is a set of beliefs about what the job is. These shifts are the real curriculum. The rest of the handbook is what the shifts look like in daily practice.
| From | To | What changes in practice |
|---|---|---|
| "Solve the ticket" | Understand why the ticket exists | They ask what the client is trying to achieve and who is affected. Half the time the ticket is a symptom of a different problem, or a request that should not be built at all. |
| "Requirements are instructions" | Requirements are hypotheses | A requirement is someone's best guess at what will fix their problem. They test the guess before building it, and they push back when the guess is wrong. |
| "I can fix it" | I can prove it | A fix without evidence is a belief. They show the reproduction, the mechanism, the test, and the metric that moved. "It works on my machine" is not proof. |
| "Code ownership" | Outcome ownership, without ego | They own the result, not the lines. If someone else's change is a better path, they take it. If their own code caused the incident, they say so first. |
| "Debugging is finding the bug" | Debugging is reducing uncertainty | Each step should cut the list of possible causes in half. They pick the check that eliminates the most possibilities, not the one that feels most likely. |
| "Wait for answers" | Unblock yourself | They find the answer in the code, the logs, the docs, or the git history. When they do need a person, they ask a precise question with a default: "I will do A unless you say otherwise by 3pm." |
| "Solo genius" | Force multiplier | Their value is measured by what the team ships, not what they ship alone. They write things down, review to teach, and build tools others can use. |
How each shift shows up in a WooCommerce agency
- Why the ticket exists. "Add a bulk export of orders" often means "finance spends 3 hours every Monday reconciling." The real fix might be a scheduled report, not an export button.
- Requirements as hypotheses. "Require phone number at checkout" is a hypothesis that phone numbers reduce failed deliveries. Ask for the data. A required field can cost 2% of conversions.
- Proof. After fixing a double-order bug, they show the log of the race being blocked, not just a week without complaints.
- Outcome ownership. The ticket is closed, the client's stock sync is still wrong. They reopen it themselves.
- Reducing uncertainty. First check: does it happen with all plugins except WooCommerce disabled? That one test rules out 60 suspects at once.
- Unblocking. Staging access is pending. They set up a local copy from the last backup and keep moving.
- Force multiplier. They spend an hour writing a WP-CLI command for the migration instead of running it by hand, so the next three migrations take minutes and anyone can run them.
Give the team a real ticket. Before touching code, each person writes: why this ticket probably exists, what the requirement assumes, what proof "done" would need, and what they would do if the client were unreachable. Compare answers. This takes 15 minutes and reveals more than a code test.
02How they solve a problem
They fix the cause, and they can prove it is the cause before they change anything.
An average developer sees a symptom and starts editing. A top-tier developer treats a bug as a claim that needs checking. What is really happening, in what order, and why? They keep what they know separate from what they assume, and they will not ship a fix they cannot explain.
- Reproduces once, in the admin, as a logged-in admin.
- Googles the error string, applies the first matching snippet.
- Adds a null check that silences the error.
- Declares "fixed" when the symptom disappears.
- Cannot explain why it broke or why the fix works.
- Reproduces under the real conditions: guest user, cached page, the specific payment gateway, the cron/AJAX context.
- Writes down assumptions and tags them
[VERIFIED]/[ASSUMED]. - Finds the earliest point in the chain where the data goes wrong, not the point where it crashes.
- Asks "what else shares this code path?" before changing it.
- Ships the fix, the reason, and a way to catch it if it comes back.
The root-cause ladder
- Observe precisely. Get the exact error, the exact request (URL, method, user role, session, cache state, gateway), and the exact time. Read the log line, not someone's summary of it.
- Reproduce deterministically. If you cannot make it happen on demand, you do not understand it yet. Use staging with a copy of production data, the same plugins, and the same cache setup.
- Locate the boundary. Narrow it down. Plugin on or off, hook priority, theme or mu-plugin, cache on or off. Find the last good state and the first bad one.
- Explain the mechanism. Write one paragraph, like: "X fires before Y because Z, so the session is not ready when the gateway reads it." If you cannot write that paragraph, keep digging.
- Fix at the right layer. Fix it where the rule was broken, not where the problem was noticed.
- Prevent recurrence. Add a log line, an alert, a test, or a guard. Something that tells you if this comes back.
Symptom: "Some customers get charged twice." Average dev adds a JS disable-on-click to the Place Order button. Top-tier dev asks: are these duplicate orders or duplicate transactions on one order? Are they from the same session? Is it the gateway retrying a timed-out request? Is Action Scheduler replaying a payment-complete job? Is a webhook arriving before the order status transition committed? The fix for each of those is completely different, and the JS fix addresses none of them.
Give a dev a bug ticket with only a customer's words ("checkout is broken"). They may ask five questions. Score them on whether the questions target context (role, gateway, device, time, cache) rather than jumping to theories.
03How they code
Boring, defensive, and written for the hostile environment the code will actually live in.
Top-tier WooCommerce code assumes the worst. It will run next to 60 other plugins, on a PHP worker that is already busy, with a cache that may be stale, on a request that may be a bot. It does not try to be clever. It tries to be correct under load and hard to misuse.
Hooks over edits, mu-plugins over theme
Business logic lives in versioned mu-plugins or a site plugin with a clear namespace. Never in functions.php, never in a core or vendor file. Survives updates and theme swaps.
Every input is hostile
Nonces on every state change, capability checks on every admin action, permission_callback on every REST route, sanitize on the way in, escape on the way out. No exceptions for "internal" endpoints.
Cost-aware by default
Knows what a hook costs and how often it fires. Never queries in a loop, never runs heavy work on init, caches expensive reads with proper invalidation, batches writes, offloads anything slow to Action Scheduler.
Uses the public API
CRUD objects and wc_get_orders(), not raw postmeta. HPOS-compatible from day one. Declares compatibility explicitly. Avoids private/internal functions that change without notice.
Safe to run twice
Webhooks, cron jobs and queue handlers are written so a replay is harmless. Guards with transients/locks or an idempotency key. Assumes at-least-once delivery.
Reads like the business rule
Function names match the domain: should_hold_for_fraud_review(), not check_order2(). A comment explains why, never what. The next dev needs no meeting to understand it.
Their mental checklist before writing a line
- Which request contexts run this code? Front-end, admin, AJAX, REST, cron, WP-CLI, Store API/Blocks checkout?
- How many times per request, and per second at peak, does this hook fire?
- What happens if this runs with a cached page, a stale object cache, or no session at all?
- What happens if it runs twice, concurrently, for the same order?
- How does it fail? Fail open, fail closed, or fail loud?
- Can it be turned off without a deploy?
- No unbounded queries (
posts_per_page => -1,get_users()without limits). - No uncached remote HTTP calls in the request path, and never without a short timeout.
- No writes to options/transients on every page load (cache invalidation storms).
- No
WC()->sessionorWC()->cartaccess in code that can run without a customer context. - No new plugin for something 40 lines of owned code can do. The judgment is not about line count. It is about security, maintenance, compatibility, ownership, lifecycle, and total cost. A plugin is a dependency you now own but do not control.
04How they maintain a codebase
They leave every file a little better than they found it, and they leave a trail.
Maintenance is where agencies win or lose margin. A codebase that only one person understands is a risk with a salary attached. Top-tier developers write for the next person, who is often themselves six months later with no memory of why.
- One place for each concern. Checkout rules in one module, shipping rules in another, integrations each in their own. No "misc" file that grows forever.
- Feature flags for anything risky. A constant, an option, or a filter that lets ops disable a behavior in seconds without a deploy.
- Decision records. A short
DECISIONS.md: "We bypass the gateway's built-in retry because it double-charged in March 2024. Do not re-enable." This is the most valuable file in the repo. - Dependency hygiene. Knows every plugin's purpose, owner, and last review date. Removes plugins that duplicate owned code. Pins versions where a silent update could break checkout.
- Deletes code. Dead hooks, commented-out blocks, "temporary" fixes from 2022. Every line that exists is a line someone has to understand.
- Upgrade discipline. Reads WooCommerce changelogs before every major, tests the checkout and order flows on staging, knows which templates are overridden and diffs them against the new versions.
- Adds a new snippet for every ticket. Fifteen snippets modify the same filter with priorities 10, 11, 12, 99, 999.
- Fixes the immediate ticket and leaves the surrounding mess.
- Knowledge lives in Slack history.
- Finds the existing module for that concern and extends it. Consolidates the priority war into one ordered pipeline.
- Refactors only what the change touches, but does it properly.
- Knowledge lives in the repo: README, DECISIONS, inline "why" comments.
Hand a dev a client site with 30+ code snippets (Code Snippets plugin, functions.php, three "custom" plugins). Task: produce an inventory. What each one does, which are dead, which conflict, and a plan to consolidate them with risk ratings. Two hours. This is what the job actually looks like.
05Entering a 100k-line unfamiliar codebase
They do not read it. They map it, starting at the entry points and following the money.
Nobody reads 100k lines. A top-tier developer builds a working picture of the system in a few hours by asking a few structural questions, then reads only the paths that matter for the task.
- Find the entry points. mu-plugins, active plugins, theme
functions.php, registered REST routes, registered AJAX actions, scheduled actions, WP-CLI commands, webhooks. This is the surface area. - Follow the money. Trace one order end to end: product page → add to cart → checkout → payment → order created → status transitions → fulfillment → email. Note every custom hook that touches this path. That's 80% of the risk in 5% of the code.
- Grep the hooks.
add_action/add_filteracross custom code, grouped by hook name. The hooks with the most listeners are the hotspots; the priorities tell you the load order. - Map the data model. What is stored where: order meta keys, custom tables, options, transients, external system IDs. Which one is the source of truth for stock, price, and customer status.
- Find the integrations. Every outbound HTTP call and every inbound webhook. Each is a place the system can hang, duplicate, or silently diverge.
- Read the git log, not the code. Last 200 commits. Hotfix clusters show where the fragile parts are. Names show who to ask.
- Run it with tracing on. Query Monitor, Xdebug, or a plain
error_log()on the path in question. Watching execution beats reading execution.
# Where is checkout being customized?
grep -rn "woocommerce_checkout\|woocommerce_after_checkout\|woocommerce_payment_complete" wp-content/ --include=*.php | grep -v "/plugins/woocommerce/"
# Who touches order status?
grep -rn "update_status\|woocommerce_order_status_" wp-content/mu-plugins wp-content/themes
# Every outbound HTTP call in custom code
grep -rn "wp_remote_\|curl_init\|new Client(" wp-content/mu-plugins wp-content/plugins/client-*
# Every scheduled job
wp action-scheduler list --status=pending --per-page=50
Clone an unfamiliar client site. In 90 minutes, the dev must produce a one-page "system map": entry points, order lifecycle hooks, integrations, sources of truth, top 3 fragile areas. Review it against what a senior already knows about the site.
06Blast radius and edge cases, spotted fast
They ask "what else uses this?" by reflex, and they carry a standard list of edge cases in their head.
Spotting edge cases quickly is not intuition. It is a memorized checklist applied without thinking. After enough WooCommerce projects, the checklist becomes reflex. Teach it directly and juniors get the reflex in months instead of years.
Blast-radius questions for any change
- Who else hooks this filter, and at what priority? Am I now running before or after them?
- Does this code path also run in REST / Store API / admin order edit / CLI / cron / subscription renewal?
- Does this change what gets cached? Page cache, object cache, WooCommerce transients, CDN?
- Does any external system read this field (ERP, Zoho, ShipStation, Klaviyo, fraud tool)?
- What happens to orders already in flight when this deploys?
- What does this do to a PHP worker under 10× traffic?
The WooCommerce edge-case list
| Domain | Edge cases top-tier devs check without being asked |
|---|---|
| Customer | Guest vs. logged in · account created at checkout · cart merged on login · session expired mid-checkout · multiple tabs |
| Cart | Empty cart · mixed physical/virtual · variation with no stock · coupon removes all items · quantity edited on checkout page · cart restored from persistent cart |
| Pricing | Tax-inclusive vs. exclusive · currency switch · role-based price · sale ending during session · rounding at 3+ decimals · negative fees |
| Checkout | Classic vs. Blocks checkout · failed payment then retry · 3DS redirect back · address autocomplete fills unexpected fields · order-pay page for pending orders |
| Payment | Gateway timeout after charge succeeded · webhook before redirect · webhook duplicate · partial refund · tokenized card removed · $0 orders |
| Orders | Status changed manually in admin · order edited after payment · stock reduced twice · renewal orders · backorders · HPOS sync mode |
| Infra | Object cache flushed at peak · cron not running · Action Scheduler backlog · CDN caching a personalized fragment · PHP worker exhaustion during a sale |
Present a one-line change ("add a 5% surcharge for COD orders"). Devs get 10 minutes to list every edge case and every other system affected. Compare lists. The gap between a junior's list and a senior's list is the curriculum.
07How they think about failure modes
They do not ask "will it work?" They ask "how will it fail, who will notice, and how fast can we undo it?"
Every production change is designed with its failure already in mind. Before deploying, top-tier developers decide what the system should do when the change goes wrong. They build the undo before they build the feature.
Fail open vs. fail closed
A fraud check that times out: block the order (lose sales) or allow it (risk fraud)? They make this a clear, written decision, never an accident of a missing else.
Fail loud
Silent failures are the most expensive kind. Every integration failure logs with context and, above a threshold, alerts a human. A stock sync that quietly stopped three weeks ago is a disaster; one that pages someone in 10 minutes is a Tuesday.
Degrade gracefully
If the recommendation engine is down, show nothing. Do not break the product page. If the shipping API is down, fall back to flat rates rather than an empty shipping section that kills checkout.
Rollback first
Before deploy: how do we turn this off in 60 seconds? Feature flag, git revert, plugin deactivate. If the answer is "restore a backup," the change is not ready.
Pre-mortem questions
- If this fails at 2 a.m. on Black Friday, what does the customer see?
- What is the worst thing this code can do to money, stock, or customer data? Can it do that twice?
- What external dependency does this add, and what's its timeout?
- Which metric would show this is broken, and are we watching it?
- What's the blast radius if the underlying assumption (e.g. "orders always have a billing email") is false for 0.1% of orders?
08Why they are much faster than average
They do not type faster. They do less, redo less, and never get stuck.
Over a month, top-tier developers are 3 to 10 times faster than average ones. Almost none of that comes from typing. It comes from work that never has to happen.
| Source of speed | What it looks like |
|---|---|
| No rework | They understand the problem before coding, so the first version is usually the shipped version. Average devs build three times. |
| Pattern recall | They have seen this kind of problem before (a gateway webhook race, for example) and know the shape of the answer. Recognition beats reasoning. |
| Fast diagnosis | They go straight to logs, Query Monitor, and bisection instead of guessing. Guessing is the slowest debugging method and the most common one. |
| Never blocked | They ask the unblocking question immediately, timebox unknowns, and have a parallel task ready. They are never "waiting." |
| Scope control | They ship the 20% that delivers 80% of the value and defer the rest explicitly. Average devs gold-plate. |
| Tooling | WP-CLI scripts, local snapshots, a personal snippet library, saved greps, browser devtools mastery. Seconds instead of minutes, hundreds of times a day. |
| Knowing the platform | They know WooCommerce has wc_get_orders() with meta queries, that Action Scheduler exists, and that Blocks checkout uses the Store API. So they do not rebuild or fight the platform. |
| AI leverage | They use AI for the parts where they can verify output instantly (boilerplate, tests, transformations, reading docs) and not for the parts where they cannot. |
Top-tier developers spend more of their time thinking and reading, and less of it typing. Their speed comes from starting later and finishing sooner.
09Choosing between options under time limits
They pick the option that is easiest to undo, and they say clearly what they are giving up.
Under deadline pressure, average developers pick the option they already know how to build. Top-tier developers pick by a fixed ranking, and they name the trade-off so the business makes the choice, not the engineer.
- Reversibility first. If two options are close, take the one you can back out of in minutes. A feature flag beats a schema change. A filter beats a template override.
- Safety over elegance. The boring, well-trodden approach (a WooCommerce hook that's been stable for 8 years) beats the clever one under deadline.
- Smallest change that fully solves the stated problem. Not the general solution. Not the platform. The problem.
- Explicit debt. If you are taking a shortcut, write the ticket for the proper version now, link it in the code comment, and tell the PM. Debt nobody has written down is the debt that hurts you later.
- Timebox the decision. 15 minutes to decide is almost always enough. Analysis past that point is usually anxiety, not diligence.
Client needs "free shipping over $100, but not for wholesale customers" by tomorrow. Options: (A) new shipping plugin, (B) custom shipping method class, (C) filter woocommerce_package_rates in an mu-plugin. Top-tier answer: C, today, because it's 30 lines, uses a stable hook, is trivially reversible and testable. They note that if the rules grow past three conditions, B becomes the right home, and file that as a follow-up.
10Handling "ASAP" and incidents
They slow down for the first five minutes so the next hour goes quickly.
Urgency is where average developers cause the second incident. The instinct to act right away is the thing that needs to be trained out. Top-tier developers follow a fixed protocol so they do not have to think about process while they think about the problem.
- Confirm impact (2 min). Is checkout down, or is one customer confused? Revenue-affecting or cosmetic? Check the last hour's order count against a normal hour. Impact determines everything after this.
- Stabilize before diagnosing. If a recent deploy correlates, revert it. If a plugin updated, roll it back. If it's a traffic spike, enable stricter caching / rate limits. Get money flowing first; understand later.
- Communicate on a cadence. One message now: what is affected, what you are doing, next update in 20 minutes. Then keep that promise even if the update is "still investigating."
- Diagnose with the ladder from section 2, but with a time limit. If you do not have a cause in 30 minutes, escalate or widen the mitigation.
- Fix minimally, deploy carefully. The incident fix is the smallest safe change. No refactoring during a fire. Verified on staging if at all possible; if not, deployed with a revert command already typed.
- Write it down. Timeline, cause, fix, detection gap, follow-ups. Blameless. Within 24 hours. This is how the team gets faster next time.
- Edit code directly on production via SFTP "just this once."
- Flush every cache at once during a traffic spike (causes a stampede).
- Run a bulk operation on live orders without a dry-run and a limit.
- Say "fixed" before verifying with a real test order on production.
- Skip telling the client because "it was only 20 minutes."
11How they test before going live
They test the way real customers behave, on a system that looks like production.
Testing on a WooCommerce site is not just a test suite (though they write those too). It is the discipline of checking the money paths under realistic conditions before a change reaches a customer.
The layers
| Layer | What | When |
|---|---|---|
| Static | PHPCS (WordPress + WooCommerce rulesets), PHPStan at a real level, PHP compatibility check | Every commit, automated |
| Unit | Pure logic: pricing rules, eligibility checks, data transforms. Fast, no WP bootstrap. | Every commit |
| Integration | WP test framework with WooCommerce loaded: does the hook fire, does the order meta save, does HPOS sync | On PR |
| Staging walk-through | Full order flows with real gateway sandbox, real plugin set, prod data snapshot, cache enabled. Guest + logged in, Classic + Blocks, desktop + mobile. | Before every deploy |
| Production smoke | One real low-value order immediately after deploy, then refund. Watch logs and error rate for 15 minutes. | After every deploy |
| Load | k6 or similar against staging for anything touching cart/checkout or adding queries to catalog pages | For performance-sensitive changes |
The staging checklist for anything touching checkout
- Guest checkout, new account checkout, returning customer checkout
- Each active payment gateway, including a declined card and a 3DS challenge
- Coupon applied, coupon removed, coupon that makes the total $0
- Cart edited on the checkout page; browser back after payment
- Order emails render correctly (customer and admin)
- Order appears correctly in every integration (fulfillment, ERP, CRM)
- Page cache and object cache enabled. Most bugs hide behind a disabled cache
- Query Monitor: no new slow queries, no N+1, no new PHP notices
If it touches cart, checkout, payment, stock, or order status, it is tested on staging with caches on and a real gateway sandbox. No exceptions for "small" changes. Small changes to checkout are how double-charges happen.
12Development and release workflow, by situation
There is no single right workflow. There is the right workflow for this project, this client, and this level of urgency.
Average developers use the same process for everything, which means it is too heavy for small fixes and too light for risky ones. Top-tier developers scale the process to the risk. The question they ask is: what is the cost if this goes wrong, and how fast could we notice and undo it?
The factors that set the workflow
| Factor | Pushes toward lighter process | Pushes toward heavier process |
|---|---|---|
| What it touches | Content, styling, admin-only screens | Cart, checkout, payment, pricing, stock, order status, customer data |
| Client size and revenue | Low order volume, downtime is an inconvenience | High volume, every minute of downtime has a dollar figure |
| Reversibility | Feature flag or git revert undoes it fully | Data migrations, schema changes, external system writes |
| Blast radius | One template, one hook, no integrations | Shared code paths, many hooks, ERP or fulfillment involved |
| Urgency | Planned work with a normal deadline | Live incident (which needs a different protocol, not a lighter one) |
| Timing | Quiet hours, mid-week | Peak season, campaign day, Friday evening |
| Team | One senior who knows the site well | Multiple devs, new to the project, or handing off to the client's team |
Four workflow tiers
Copy, CSS, admin-only tweaks
Branch, self-review, deploy to staging, quick visual check, deploy to prod. Same day. No change ticket needed beyond the commit message. Still never edited directly on production.
Features off the money path
Ticket with acceptance criteria, branch, PR with peer review, staging walk-through, deploy in a normal window, smoke test after. One to three days.
Anything on cart, checkout, payment, stock, orders
Written spec with edge cases, PR reviewed by a senior, full staging checklist with caches on and gateway sandbox, deploy in a low-traffic window with a named rollback, real test order on production, monitoring for 24 hours. Client informed before and after.
Data migrations, schema changes, external system writes
Everything in Tier 3, plus: dry run on a production data copy, row counts before and after, a backup taken immediately before, a written data rollback plan, and a go/no-go checkpoint with the client.
Adjusting for the client
- Small client, low volume. Tier 1 and 2 cover most work. Do not make them pay for enterprise process on a $50/day store. But their checkout is still Tier 3, because it is 100% of their revenue.
- High-traffic client. Nothing is Tier 1 if it touches a template that renders on every page. A CSS change that breaks the cache key is a Tier 3 event. Deploy windows and monitoring are not optional.
- Client with their own team. Heavier documentation, lighter ceremony. They will maintain it after you, so the PR description and the DECISIONS entry matter more than the meeting.
- Retainer vs. project. On a retainer, invest in tooling that makes future deploys cheaper (staging refresh scripts, deploy pipelines). On a fixed project, keep it simple and document the handover.
Branching and environments
- Trunk-based with short-lived feature branches for most agency work. Long-lived release branches only for clients with formal release calendars.
- Three environments minimum for Tier 3 clients: local, staging (refreshed from production regularly, with customer data masked), production.
- Staging must have the same plugin set, the same cache stack, and the gateway in sandbox mode. A staging site that does not look like production is a false sense of safety.
- Deploys are scripted and repeatable. If a deploy requires remembering steps, it will be done wrong under pressure.
- Every deploy has a tag. Every tag can be redeployed in under five minutes.
An urgent Tier 3 change is still a Tier 3 change. Urgency compresses the calendar, not the checklist. If the checklist cannot be completed in the time available, the honest answer is a smaller change or a temporary mitigation, not a full change with steps skipped.
Give the team six real tickets from different clients. For each, they assign a tier and justify it in one sentence. Discuss the disagreements. The disagreements are where the team's risk model is unclear.
13Backups and disaster recovery
A backup you have never restored is a hope, not a backup. And code rollback, data rollback, and external-system rollback are three different problems.
Most agencies believe they have backups because the host says so. Top-tier developers know what is backed up, how often, how long a restore takes, and what would be lost. They have tested it. And they understand that "roll back" means three different things depending on what went wrong.
Three kinds of rollback
Easy, fast, and usually complete
Redeploy the previous tag, revert the commit, or flip the feature flag. Takes minutes. The catch: code rollback does nothing about what the bad code already did to the data while it was live.
Hard, slow, and rarely complete
Restoring last night's database loses every order placed since then. You almost never restore the whole database on a live store. Instead you need targeted repair: identify affected rows, fix them with a script, verify counts. This is why you take a backup immediately before any risky change, not just nightly.
Often impossible
The bad code sent 300 orders to ShipStation, charged cards through the gateway, pushed stock levels to the ERP, or triggered Klaviyo emails. None of that lives in your database. Rollback means contacting each system, running their bulk APIs, or manually undoing. Sometimes it cannot be undone at all. This is why external writes get the heaviest process.
A pricing bug ships at 9am and sets all sale prices to zero. Code rollback at 9:20 stops new damage. But 40 orders were placed at $0, stock was reduced, order confirmation emails went out, and the fulfillment system already has them. Reverting the code fixes none of that. The real recovery is: cancel or contact 40 customers, restore stock, pull the orders from fulfillment, and explain to the client. The 20-minute code fix was the easy part.
What a real backup strategy covers
| Question | Top-tier answer |
|---|---|
| What is backed up? | Database, uploads, code (in git anyway), and the configuration that is not in either: server config, cron entries, environment variables, DNS, CDN rules, gateway settings. |
| How often? | Database at least hourly on high-volume stores, plus a snapshot before every Tier 3 or 4 deploy. Files daily. Off-site copy, separate from the host. |
| How far back? | Enough to catch slow damage. A stock sync that has been quietly wrong for two weeks needs a two-week-old copy to compare against. |
| How long to restore? | Known, because it has been timed. On a 20GB database, a full restore can take hours. That number decides whether restore is even an option during an incident. |
| Who can do it? | At least two people, with a written runbook, without needing the person who set it up. |
| When was it last tested? | A date within the last quarter, with a note of what went wrong during the test. Something always does. |
Disaster recovery on a WooCommerce store
- Define the scenarios. Host down, database corrupted, site hacked, bad deploy, bad plugin update, domain or DNS compromised, key staff unavailable. Each has a different runbook.
- Know the two numbers. How much data can we afford to lose (hours of orders)? How long can we afford to be down? The client decides these, and they set the backup frequency and the hosting budget.
- Plan for partial recovery. Restoring to a clean copy from four hours ago and replaying the orders from the gateway's transaction log is often better than a full restore. Payment gateways and email logs are your secondary record of what happened.
- Keep a maintenance mode that works. A static page that tells customers the truth, served from the CDN, that can be turned on in one minute. Better than a broken checkout taking money it cannot record.
- Practice. Once a quarter, restore production to a fresh staging environment from backups alone. Time it. Fix what broke. This is also how you refresh staging.
- Fresh database snapshot, verified by size and a row count on the affected tables.
- Written list of every external system the change writes to, and how each write would be undone.
- Dry run on a copy of production data, with before and after counts.
- The rollback command or script written and tested before the forward change runs.
- Someone watching order flow for the first hour after.
Scenario: a bad stock sync ran for six hours and wrote wrong quantities to 800 products, and the ERP pulled those numbers. Team writes the recovery plan in 30 minutes: what to roll back, in what order, what cannot be rolled back, and what to tell the client. Then, separately: restore last night's backup to a staging server and time it.
14Realistic estimation
Timeboxing controls how long you spend. Estimation predicts how long it will take. Top-tier developers do both, and they know they are different skills.
Timeboxing is a discipline: "I will spend two hours on this, then reassess." Estimation is a forecast: "This will take about two days." The first protects you from sinking time. The second is what the client is asking for when they say "how long?" Average developers estimate the typing. Top-tier developers estimate the whole job.
Why developer estimates are wrong
- They estimate the happy path and forget the edge cases, the review round, the staging test, and the deploy.
- They estimate for themselves on a good day, with full focus and no interruptions.
- They do not count the unknowns: an undocumented plugin, a client who takes two days to answer, a gateway sandbox that behaves differently from live.
- They give one number when the honest answer is a range.
How top-tier developers estimate
- Break it down until each piece is under half a day. Anything larger is hiding an unknown. If you cannot break it down, that is the estimate's biggest risk, and you say so.
- Count the whole lifecycle. Understanding, building, edge cases, review, staging test, deploy, monitoring, client communication. On a Tier 3 change, typing is often less than a third of the total.
- Name the unknowns and price them. "Two days if the gateway supports partial capture, four if we have to work around it. I will know by tomorrow noon." An estimate with a named risk is worth more than a confident number.
- Give a range, and a confidence. "Three to five days, likely four." Ranges tell the client how sure you are. A single number hides that.
- Check against history. How long did the last three similar tasks actually take? Most developers are consistently off by the same factor. Learn yours.
- Re-estimate at the first checkpoint. After the first few hours you know far more. Update the estimate then, not on the due date.
Client asks: "How long to add a gift wrap option at checkout?" Average answer: "Half a day." Top-tier answer: "The field and the fee is two hours. But it needs to show on the order, in the emails, in the packing slip, and go to fulfillment as a line item so the warehouse sees it. It needs to work in Blocks checkout, which the site uses. And it needs a test on staging with the real fulfillment connection. Two to three days. The main unknown is whether ShipStation accepts custom line items the way we need. I can confirm that in an hour before you decide."
When timeboxing and estimation meet
- Use a timebox to produce the estimate: "Give me two hours to investigate, then I will give you a real number."
- Use a timebox inside the work to catch overruns early: if a half-day piece hits a full day, stop and report, do not push through silently.
- Never let a timebox become a deadline promise. "I will spend a day on it" and "it will be done in a day" are different sentences, and clients hear the second one.
Every sprint, each developer writes their estimate and range before starting a ticket, then records the actual. After a month, review the ratio per person. Most people are off by a consistent factor, often between 1.5x and 3x, in the same direction each time. Once they know their factor, their estimates become useful. Keep doing this permanently.
15The quality model
"Working" is one dimension out of eight. Top-tier developers check all eight before calling something done.
Ask an average developer if their feature is finished and they will tell you it works. Ask a top-tier developer and they will tell you how it scores on each of these. This is the agency's definition of quality, and it is the checklist for every review.
Security
Nonces, capabilities, permission callbacks, sanitizing, escaping, no exposed endpoints, no secrets in code, no privilege paths a customer can reach. Assumes hostile traffic.
Performance
No new queries per page load, no uncached remote calls, no unbounded loops, heavy work off the request. Measured with Query Monitor and under load, not assumed.
Accessibility
Keyboard reachable, labeled form fields, visible focus, sensible contrast, screen reader announces errors. A checkout that a customer cannot complete with a keyboard is a lost sale and a legal risk.
SEO
Correct status codes, no accidental noindex, canonical URLs intact, structured data still valid, page speed not regressed, no duplicate URLs created by filters or parameters.
Compatibility
Works with the current WordPress and WooCommerce, HPOS, Blocks checkout, the site's theme and plugin set, and the PHP version. Does not break on the next WooCommerce minor release.
Reliability
Handles failure of every dependency it touches. Idempotent where it must be. Logs what matters. Can be turned off without a deploy. Does not depend on cron running on time.
Maintainability
Lives in the right module, follows the codebase's conventions, explains its reasons, has no hidden coupling, and can be understood by the next developer without a meeting.
Business correctness
Does what the business actually needs, not just what the ticket said. The tax is right. The discount stacks the way finance expects. The report matches what the accountant sees. Verified with the person who owns the outcome.
Using the model
- In planning: which dimensions does this change touch? A checkout change touches all eight. An admin report touches maybe four. That decides the workflow tier.
- In review: the reviewer walks the eight dimensions, not just "does the code look right." Most missed bugs live in a dimension the reviewer did not think about.
- In done criteria: "done" means each relevant dimension has evidence, not a feeling. A screenshot of Query Monitor, a keyboard walk-through, a test order with the right tax.
- In trade-offs: when time is short, say which dimension you are trading. "We are shipping without the accessibility pass; ticket filed for Tuesday." Never trade security or business correctness.
Take a finished PR from last week. Each reviewer scores it 1 to 5 on all eight dimensions, independently, with one sentence of evidence per score. Compare. The dimensions where scores differ most are the ones the team does not have a shared standard for yet.
16How they know a problem actually matters to the business
They turn every ticket into revenue, cost, or risk before deciding how much attention it deserves.
Average developers treat every ticket as the same kind of engineering problem. Top-tier developers know that a slightly ugly admin screen and a 2% drop in checkout conversions are not the same kind of problem, and they spend their attention accordingly, even when the client is loudest about the admin screen.
The three questions
- Does it touch money? Cart, checkout, payment, pricing, stock, refunds. If yes, it's P1 until proven otherwise.
- How many customers, how often? "One customer once" and "every mobile Safari user" are different tickets with the same description.
- What's the cost of being wrong? A wrong shipping label costs $8. A wrong tax calculation costs an audit. A leaked customer list costs the client's business.
They also know the client's numbers: average order value, daily order count, conversion rate, peak hours, top products. With those numbers, "checkout is 400ms slower" becomes "roughly X fewer orders per day," and priorities set themselves.
Two tickets arrive together: "Product filter looks off on tablet" (from the CEO) and "Occasionally the order confirmation email doesn't arrive" (from support). The top-tier dev checks: how many orders per day, how often is "occasionally," what the support cost of each one is, and could the same root cause be dropping other transactional emails? They do the email first and tell the CEO why in one sentence.
17How they communicate
Short, structured, honest about what they do not know, and always with a recommendation.
- "I think it might be the caching plugin or maybe the theme, still looking."
- Reports activity ("worked on the checkout bug").
- Asks open questions ("what do you want me to do?").
- Surprises people with problems at the deadline.
- "Cause confirmed: the fraud plugin's webhook fires before the order commits. Fix: defer via Action Scheduler. Risk: low. On staging by 3pm, prod tonight after review."
- Reports outcomes and next steps.
- Asks closed questions with a recommendation ("Option A or B? I recommend A because…").
- Raises risk the moment it appears, with options.
Templates the team should use
Done: what shipped. Found: anything surprising. Next: what happens next and when. Need: decisions or access you are waiting on (or "nothing").
Problem in one sentence · Options ranked (2–3 max) · Recommendation + why · Risk and rollback · Effort.
PR description answers: What changed · Why · What contexts it runs in · What could break · How it was tested · How to turn it off.
They also write for non-engineers. The client hears "orders will no longer be held when the fraud service is slow," not "I deferred the webhook via a queue."
18Teamwork and how they add value
They multiply the team, not just their own output.
- Their reviews teach. Comments explain the principle ("this runs on every page load; move it to a transient with a 10-minute TTL because…"), not just the fix. Juniors level up from reading them.
- They make themselves unnecessary. Documentation, runbooks, feature flags, and clear code mean anyone on the team can operate their work. Bus factor of one is a failure, not a flex.
- They disagree early and commit fully. Push back in planning with reasons; once decided, execute without relitigating.
- They own outcomes, not tickets. If the ticket is closed and the client still has the problem, the work is not done.
- They notice the unasked question. "You asked for a CSV export. Did you know the ERP already has an API that would make this unnecessary?" Sometimes the highest-value work is the work you talk the client out of.
- They reduce the team's risk. Standards, checklists, automation, and the occasional "please don't merge that on a Friday."
How they work on other people's code
- They assume the previous developer had a reason, and they look for it (git blame, comments, tickets) before removing anything.
- They match the existing style even when they would have done it differently. Consistency beats preference.
- They do not rewrite what they can extend. A rewrite is a business decision, not a personal preference.
- They add tests around legacy code before changing it. The tests describe current behavior, so the change shows up as a clear difference in behavior.
- They leave the file better: a clarifying comment, a removed dead branch, a renamed variable. Small, safe, every time.
Training the team in the age of AI
When any developer can generate a working plugin in a minute, the skills that always set people apart become the only thing that sets them apart.
19What AI actually changed
It made producing code almost free. It did nothing to the cost of being wrong.
A junior developer's value used to be "can produce working code." That value is now close to zero. An AI produces working-looking code faster, and the client can prompt it too. What is still valuable, and has become more valuable, is everything in Part I: knowing what to build, knowing whether what was built is correct, knowing what it will do to the rest of the system, and being accountable when it breaks.
Cheap now
Boilerplate, CRUD, hook wiring, template overrides, REST endpoints, documentation drafts, test scaffolding, translating between formats, reading unfamiliar APIs.
Still expensive
Correct problem definition. Knowing the platform's real behavior (not the docs). Judging blast radius. Verifying under load and concurrency. Deciding what not to build. Owning the outcome.
Newly dangerous
Plausible code that's subtly wrong: non-HPOS meta access, missing nonces, queries on every page load, hooks in the wrong context, race conditions that only show at scale. AI produces these fluently and confidently.
Newly essential
The skill of verifying: reading generated code faster than you could write it, and knowing exactly which claims to test. The reviewer is now both the bottleneck and the safety layer.
The new failure mode to train against
A developer who accepts generated code they cannot explain has not become faster. They have become a bigger risk. The rule for the agency: if you cannot explain every line in a PR to a senior without looking, you cannot submit it. AI is allowed for everything; understanding is required for everything.
How top-tier devs use AI
- They write the spec, constraints, and edge cases first, the same discipline as before, and hand that to the AI. Vague prompt, vague code.
- They use it for the parts they can verify in seconds and do the parts they cannot verify by hand.
- They ask it to attack its own output: "list the ways this fails under concurrent requests," "what assumptions does this make about the order object?"
- They use it to read: summarize a plugin's architecture, find every hook that touches X, explain a legacy function.
- They never let it make the architectural decision. They make it, then use AI to execute.
- They keep project context in the repo (a
CLAUDE.md/AGENTS.mdwith stack, conventions, and forbidden patterns) so generated code starts from the team's standards, not the internet's average.
20The training program
A 14-week structure you can repeat. Every week has the same theme: judgment, verification, ownership.
Principles for the program
- Train on real client code, anonymized. Toy projects do not have plugin conflicts, cache layers, or five years of snippets.
- Grade explanations, not output. A dev who ships a correct fix and cannot explain it scores lower than one who ships a partial fix with a full root-cause write-up.
- Make seniors teach. The fastest way to make a senior top-tier is to have them articulate their instincts for a junior. Pair every drill.
- Timebox everything. Speed under a limit is the skill. Unlimited time produces gold-plating.
- Every drill ends in writing. A one-page document: what, why, risk, what they would do with more time. Communication is trained every week, not as a separate module.
| Week | Theme | Core activity | Output graded |
|---|---|---|---|
| 0 | Mindset | The seven shifts, applied to five real tickets before any code is written | Written answers per ticket |
| 1 | Root cause | Assumption tagging on 3 real bug tickets; reproduce, bisect, explain | Mechanism paragraph for each |
| 2 | Reading code | Map an unfamiliar client site in 90 minutes | One-page system map |
| 3 | Edge cases | Edge-case listing drills on 5 one-line changes | List compared to senior's |
| 4 | Blast radius | For 3 PRs, document every context and consumer affected | Impact analysis |
| 5 | WooCommerce internals | Trace one order through core: session → cart → checkout → order → payment → status. Read the actual source. | Annotated flow diagram |
| 6 | Performance | Profile a slow catalog page with Query Monitor; find and fix N+1 and uncached calls under load | Before/after with numbers |
| 7 | Security | Audit 5 custom AJAX/REST endpoints from real projects; find missing checks | Findings + fixes |
| 8 | Failure modes | Pre-mortem 3 planned features; design fail-open/closed, alerts, rollback | Pre-mortem docs |
| 9 | AI verification | Review 10 AI-generated WooCommerce snippets; find the subtle bugs (there are almost always some) | Bug list with severity |
| 10 | Incidents | Simulated incident on staging (broken deploy during a "sale"); run the protocol | Timeline + post-mortem |
| 11 | Decisions | 3 scenarios with deadlines; choose between options, name trade-offs, write the proposal | Proposal doc |
| 12 | Release and recovery | Tier six real tickets; write a recovery plan for a bad data sync; restore a backup to staging and time it | Tier justifications + recovery plan + restore time |
| 13 | Estimation and quality | Estimate with ranges on real tickets and record actuals; score three PRs on all eight quality dimensions | Estimate log + scored reviews |
| 14 | Ownership | Own a real small client change end to end: spec, build, test, deploy, monitor, report | The whole thing |
Ongoing rituals after the 14 weeks
- Weekly bug autopsy (30 min): one real bug from the week, walked from symptom to mechanism by whoever fixed it. Everyone learns the pattern.
- Monthly "explain this AI code" session: a generated PR, reviewed live. Find the flaw.
- Every incident gets a blameless post-mortem shared with the whole team within 24 hours.
- Review comments must teach. Enforce it. "Fix this" is not a review comment.
- Maintain the team's
DECISIONS.mdand forbidden-patterns list. Every incident adds a line.
21Review rubric and levels
What "junior," "senior," and "top-tier" actually mean, so promotions are based on evidence, not feelings.
| Dimension | Junior | Senior | Top tier |
|---|---|---|---|
| Problem solving | Fixes the symptom with guidance | Finds root cause independently | Finds root cause, prevents the class of bug, and improves detection |
| Code | Works; needs review for security/perf | Secure, performant, follows standards | Also idempotent, observable, flag-controlled, and readable by anyone |
| System understanding | Knows the feature | Knows the site | Knows the site, its integrations, its infra, and its business numbers |
| Edge cases | Handles the ones pointed out | Finds most before review | Finds them in the spec phase and changes the spec |
| Risk | Unaware of blast radius | Assesses it when asked | Assesses it unprompted; designs the rollback first |
| Speed | Slow, with rework | Steady | Fast because of no rework and no blocking |
| Communication | Reports activity | Reports outcomes | Reports outcomes with options, recommendations, and risk, to any audience |
| AI usage | Copies output | Reviews output | Directs, constrains, and verifies output; teaches others to |
| Estimation | Guesses the typing time | Ranges with the lifecycle counted | Ranges with named unknowns, re-estimated at checkpoints, calibrated against their own history |
| Release discipline | Same process for everything | Matches process to risk when reminded | Tiers every change unprompted; knows the three kinds of rollback and plans data recovery before deploying |
| Quality | Checks that it works | Checks security and performance | Walks all eight dimensions with evidence and names any trade-off out loud |
| Team | Consumes help | Gives help | Raises the team's floor: standards, docs, reviews that teach |
PR review checklist (apply to every PR, human or AI-written)
- Can the author explain every line without looking?
- Which request contexts execute this? Is it guarded for the ones it should not run in?
- Nonce, capability, permission_callback, sanitize, escape: all present where needed?
- Any query or remote call that runs per page load or per loop iteration?
- HPOS-safe? Blocks-checkout-safe? Works with caches on?
- Safe to run twice? Safe under concurrent requests for the same order?
- What breaks if this is wrong, and how is it turned off?
- Tested how? Evidence in the PR?
- Which of the eight quality dimensions does this touch, and is there evidence for each?
- If this needs undoing, what is the code, data, and external-system rollback?
- Would a new hire understand this file in six months?
22Drill library
Short exercises you can repeat. Rotate them. Score them. The scores show you where the team needs work.
Five questions
Bug ticket in customer language. Dev gets five questions before any theory. Score question quality.
Edge-case sprint
One-line feature request. List every edge case and affected system. Compare against a senior's list.
Find the flaw
AI-generated WooCommerce snippet that looks right. Find the security, performance, or concurrency bug. There is almost always one.
System map
Unfamiliar client repo. Produce entry points, order flow hooks, integrations, top fragile areas.
Pre-mortem
Planned feature. Write how it fails at peak, who notices, and the 60-second rollback.
Priority war
Filter with 8 competing callbacks at scattered priorities. Consolidate into one predictable pipeline without changing behavior.
Staged incident
Senior breaks staging in a realistic way during a mock sale. Junior runs the incident protocol. Grade on process, not just fix time.
Translate it
Take a technical PR description and rewrite it for the client. Then for the CEO. Then as a one-line Slack update.
Cost this hook
Given a callback and the hook it's on, estimate executions per second at peak and DB cost. Then measure with Query Monitor. Compare.
Estimate it
A real ticket. Write the breakdown, the range, the named unknown, and the confidence. Record the actual later. Track the ratio per person.
Three rollbacks
A bad deploy ran for hours. Write what code rollback fixes, what data needs repair, and which external systems were written to and how to undo each.
Tier it
Six tickets, six tiers, one sentence of justification each. Discuss where the team disagrees.
Eight dimensions
Score a finished PR on the quality model with one line of evidence per dimension. Compare scores across reviewers.
Pick under pressure
Scenario + deadline + 3 options. Choose, rank, name the trade-off, file the debt ticket. Defend it in two sentences.
Every drill trains the same habit: asking "what is actually true, what could go wrong, and who is affected" before acting. That habit is what the agency sells. Code is the by-product.