Engineering Culture

Toil: When to Automate and When to Grind

Google's SRE teams cap manual, repetitive 'toil' at 50% of their time. Past that, the operational work eats the engineering that would reduce it.

RE

Roberto Espinoza

CEO, Ruzora

July 26, 20268 min read

Every engineering team accumulates a kind of work that feels productive but builds nothing: the manual deploy step, the weekly report someone assembles by hand, the recurring incident everyone knows how to clear but nobody has fixed. Google's SRE teams named it toil, and set a hard rule about it that's worth stealing: keep it under half your time, or it will quietly eat all of it.

Key Takeaways

  • Toil is work that's manual, repetitive, automatable, tactical, and without enduring value, scaling linearly with the service (Google SRE).
  • Google caps toil at under 50% of each engineer's time, reserving the rest for engineering that reduces future toil (Google SRE).
  • Left unchecked, toil expands to fill 100% of everyone's time (eliminating toil).
  • The threshold turns "should we automate this?" into a budget question, not a vibe.

What Toil Is (and Isn't)

Google's definition is precise (Google SRE on toil). Toil is work tied to running a service that is manual, repetitive, automatable, tactical (reactive rather than strategic), devoid of enduring value (the service is no better after you do it than before), and scales linearly as the service grows. That last property is the killer: as your usage doubles, linear toil doubles too, so it grows without bound unless you automate it away.

Note what toil is not. Overhead like meetings and email is plain overhead rather than toil. And genuinely hard, one-off problem-solving isn't toil even when it's unpleasant, because it has enduring value. Toil is specifically the repetitive, automatable operational grind, the stuff a script could do.

The 50% Rule

The SRE guardrail is that operational work (toil) should stay below 50% of each engineer's time, leaving at least half for engineering projects that either reduce future toil or add real capability (eliminating toil). This is treated as a management guardrail rather than an aspiration, because toil has a nasty dynamic: it tends to expand to fill whatever time you give it, and can quietly reach 100% of a team's hours. Once it does, there's no time left to automate anything, so the team is trapped, permanently busy with work that builds nothing and grows every quarter.

If toil crosses 50%, the SRE book treats it as a management problem needing intervention: add people, redirect engineering effort at the worst toil, or push back on whatever keeps generating it.

Real engineeringToil
Builds enduring valueLeaves the service unchanged
Scales sub-linearly (automate once)Scales linearly with growth
StrategicTactical, reactive
Should be over 50%Should be under 50%

A Concrete Version

A three-person team spends more and more of each week on manual releases, hand-run data fixes, and clearing the same recurring alert. Nobody automates it, because they're too busy doing it, the exact trap. Six months later they're at roughly 80% toil, shipping almost no features, and burning out, while every uptick in customers makes it worse because the toil scales with usage. The fix is to treat the 50% line as a hard budget: block engineering time to automate the biggest toil source even though it hurts short-term, because that one script buys back hours every week forever.

The Honest Counterpoint

The 50% number is Google's, at Google's scale, and it isn't sacred for a five-person startup. Early on, a scrappy team runs on plenty of manual work because automating everything prematurely is its own waste, you don't script a process you're still figuring out or that runs twice. There's also a real threshold calculation: automation pays off only when the time saved over the automation's life exceeds the time to build and maintain it, so some rare, cheap toil is correctly left manual (the automation tradeoff). The durable lesson is to measure toil, cap it deliberately, and automate the repetitive, linearly-scaling grind before it consumes the team, rather than the specific number.

What This Means for Teams

Toil is a slow, quiet way for a team to become unable to improve itself, and it's especially dangerous for lean startups where one or two people carry all the operations. The reflex to notice "we do this by hand every week and it's growing" and to protect time for automating it is a senior trait, and it connects to why 100% utilization backfires: a team with zero slack has no room to automate the toil that's drowning it. When toil is eating a team, the honest fixes are the SRE ones, more capacity or dedicated automation time, rather than working harder. See available engineers.

Frequently Asked Questions

What is toil in engineering?

Google SRE's term for work that is manual, repetitive, automatable, tactical, without enduring value, and that scales linearly as the service grows. The manual deploy, the hand-built report, the recurring alert you keep clearing.

What's the 50% rule?

Google keeps toil under 50% of each engineer's time, reserving the rest for engineering that reduces future toil or adds capability. Past 50%, it's treated as a management problem needing intervention.

Why is toil so dangerous?

Because it scales linearly with growth and expands to fill available time. Left unchecked it can reach 100% of a team's hours, at which point there's no time left to automate it away, trapping the team.

Should a small startup enforce 50%?

Not rigidly. Early teams run on manual work, and automating prematurely is wasteful. The durable lesson is to measure toil, cap it deliberately, and automate the repetitive, growing grind before it consumes the team.

The Bottom Line

Toil, the manual, repetitive, automatable work that scales with your service, expands to fill whatever time you give it, until there's no time left to automate it away. Google caps it at half of each engineer's time for exactly that reason. Measure your toil, protect time to automate the worst of it, and treat crossing the line as a staffing or scope problem rather than a call to grind harder.

Roberto Espinoza is CEO of Ruzora, which helps US startups hire pre-vetted senior LATAM engineers, with a vetted shortlist in 72 hours. See available engineers.

RE

Roberto Espinoza

CEO, Ruzora

Roberto is the founder and CEO of Ruzora. He works directly with US startup founders and CTOs on staff-augmentation and software-factory engagements, and personally reviews senior engineer placements.

AI-vetted engineers, ready now

Your next senior engineer is already vetted and waiting.

It starts with a single call. 72 hours later, you're reviewing scored candidates who already match your stack and culture.