"Improve the product" is not a plan
Look at your team's current quarter. Somewhere in it is an objective that sounds like improve onboarding, or make the platform more reliable, or be more data-driven. Everyone nodded. Nobody disagreed. And in thirteen weeks somebody will ask whether you did it, and the honest answer will be that it depends who you ask — which means the sentence was never a commitment at all, just a direction of travel with a deadline attached.
The thing that makes an OKR work is not the Objective. It is the two or three Key Results underneath it, because a Key Result is the only part that can be wrong. Reduce median time-to-first-value from 9 days to 3 is a claim about the future that can fail. Improve onboarding cannot fail; it can only be discussed. The discipline OKRs impose is not ambition. It is the requirement to state, in advance and in writing, what would count as having succeeded.
That is a decision-record discipline rather than a goal-setting one, and reading OKRs that way explains both why they help and why so many rollouts quietly die. This piece covers where they actually came from, what the research supports and what it does not, the strongest published case against goal-setting, and the reframe that makes the practice survive contact with a real quarter.
Where OKRs came from — and the part that is contested
Andy Grove introduced the practice at Intel around 1971, where it was first called iMBO — Intel Management by Objectives. He set it out in High Output Management in 1983. The Intel version already had the features people now treat as definitional: goals set quarterly rather than annually, negotiated between manager and employee rather than issued downward, and graded numerically on a 0 to 1.0 scale.
John Doerr was an engineer at Intel in the 1970s and among the first people taught the system in Intel's own training. In 1999 he presented it to a young Google, and his 2018 book Measure What Matters carried it to a general audience. Doerr has been consistent about the credit: Grove was the inventor; he was the messenger.
The tidy origin story has a real objection against it, and it is worth stating because it changes how much authority the framework deserves. Critics — the Balanced Scorecard Institute among them — point out that OKR and MBO are based on the same principles and follow roughly the same process, MBO being Peter Drucker's, from the 1950s. On this reading Grove did not invent a management system; he fixed a broken implementation of an existing one. The Institute concedes the achievement plainly enough: Grove was smart enough to see why the Management by Objectives implementation at Intel needed improvement. That is a genuine contribution. It is just not a new invention, and a framework inherits the evidence of its parent rather than starting fresh.
What the research actually supports
OKRs themselves have almost no direct empirical literature. What they have is a parent theory with a great deal of it. Goal-setting theory began with Edwin Locke in 1968 and was developed with Gary Latham through the 1970s and 1980s; the pair summarised it in American Psychologist in 2002 under the title Building a Practically Useful Theory of Goal Setting and Task Motivation: A 35-Year Odyssey.
The core finding is stable and unusually well replicated: specific, difficult goals produce higher performance than urging people to do their best. Vagueness is not a neutral choice — it measurably underperforms. That single result is the strongest thing anyone can say in defence of writing Key Results down, and it is a strong thing.
But the theory is not the slogan. Locke and Latham are explicit that the effect is conditional, and the conditions are where organisational practice goes wrong.
The conditions everyone drops
Goal-setting theory attaches moderators — conditions under which specific difficult goals help, and outside which they do not. Read them next to a typical OKR rollout and the failure modes name themselves.
- ✓Commitment. The goal has to be genuinely accepted, not merely received. A Key Result assigned to a team that thinks it is the wrong number is not a difficult goal; it is a compliance exercise, and the theory does not predict it will work.
- ✓Feedback. People need information about progress. A number reviewed once at quarter-end is not feedback — it is a verdict. The Intel version graded quarterly for a reason.
- ✓Ability. Difficult must mean difficult, not impossible. Where the skill or the capacity is absent, a stretch target produces neither performance nor learning.
- ✓Task complexity. On genuinely complex work, goal intensity can arrive too early — the theory allows for a learning phase first, which is precisely the case where quarterly targets sit worst.
- ✓Clarity. Specific and measurable, not aspirational. This is the one that improve the product fails, and it fails at the first hurdle rather than the last.
Take one Key Result from your current quarter and ask which of the five conditions it actually satisfies. Most survive clarity and fail commitment — the number is specific, and nobody who has to hit it believes in it. That is not an OKR problem; it is an unrecorded disagreement, and it will surface as a missed target three months later instead of as an argument now.
The strongest case against goals
In 2009 four researchers — Lisa Ordóñez, Maurice Schweitzer, Adam Galinsky and Max Bazerman — published Goals Gone Wild: The Systematic Side Effects of Overprescribing Goal Setting in Academy of Management Perspectives. Their argument is that the benefits of goal setting have been overstated while the harms have been largely ignored, and that goals should be treated like a potent medication: prescribed carefully and monitored closely. The side effects they document are specific.
The honest thing to add — and most summaries of this paper omit it — is that Locke and Latham did not accept the charge. They published a rebuttal in the same volume of the same journal, arguing that the critics had abandoned good scholarship, and the exchange remains live rather than settled. Anyone telling you the science has decided against goals is overstating it in exactly the direction the original paper warns about.
- ✓Narrowed focus. Attention concentrates on the goal and drains from everything not measured by it.
- ✓Distorted risk preferences. A target changes how much risk people are willing to take, usually in the direction of the target rather than of the business.
- ✓Unethical behaviour. Goals raise the incidence of it — the best-documented and most uncomfortable of the findings.
- ✓Inhibited learning. Performance goals compete with learning; on unfamiliar work, hitting the number can cost the understanding.
- ✓Corroded culture and reduced intrinsic motivation. Extrinsic targets crowd out the reasons people were doing the work anyway.
The reframe: a Key Result is a decision record
Both sides of that dispute are arguing about goals as incentives — things that pull behaviour. That is the use with the side effects, and it is the use that requires all five moderators to hold. There is a second use that neither side really contests, and it is the one worth keeping.
A Key Result written before the work begins is a record of a decision: this is what we agreed success would look like, at the moment we agreed it, before anyone knew whether we would get there. That artifact is valuable even if you never grade it. It fixes the definition of good while the definition is still arguable, rather than after results are in and everyone's memory has helpfully adjusted. It is the same move that a written pre-read makes against slides, and the same move a decision record makes against a meeting summary.
Read that way, the familiar OKR failure is not a goal-setting failure at all. It is the failure the whole decision-record literature describes: the reasoning that produced the number is nowhere. Six months on you can see that the team committed to three days and delivered nine, and you cannot see which assumption was wrong, who doubted it at the time, or what you would now do differently — the same gap that turns a closed action item into an unanswerable question.
What to change on Monday
None of this requires abandoning the framework. It requires treating the writing as the point and the grading as secondary.
- 1Write the disagreement down next to the number. When a Key Result is set, record who thought it was wrong and why. Commitment is a moderator; an unrecorded objection is the moderator failing silently.
- 2Record the assumption each number rests on. Three days assumes the new import pipeline ships in week two. When the target is missed, you learn something instead of arguing about effort.
- 3Separate learning quarters from performance quarters. Task complexity is a moderator too. On genuinely new work, set a learning goal and say so, rather than setting a number nobody can honestly forecast.
- 4Review the reasoning, not just the score. A 0.6 against a sound assumption is a better quarter than a 1.0 against a target that was set low because someone was being measured on it.
- 5Keep the record after the quarter closes. The score expires; the definition of good and the argument behind it are what the next planning round needs, and they are exactly what most OKR tools discard.

