This essay builds on Charles Kent’s in Agent Autopsies, where he argues that AI has broken the metrics organisations use to judge employee performance. Metrics based on volume, such as support tickets resolved or pull requests shipped, once served as reliable proxies for effort and skill. Those output-based measures no longer provide the same signal now that AI enables people to produce competent work much more quickly.
Charles argues that performance management in 2026 must evolve beyond traditional output-based measures. Alongside outcomes and quality, organisations should also assess how effectively employees use AI to create value.
That argument resonated with me, because I spent years in HR and Talent Acquisition making hiring decisions, assessing capability and evaluating performance. Much of that work relied on interpreting signals that, until recently, were relatively stable. AI is changing that. The assumptions many organisations have relied on to judge performance are becoming less dependable as people and AI increasingly co-produce work.
That raises the question I want to explore in this essay: if AI is changing how work is produced, how should organisations evaluate the distinctly human contribution?
—
For decades, organisations have relied on a simple assumption: if someone consistently produced high-quality work, they were probably thinking well.
It was never a perfect measure of performance, but it was a practical one. Managers cannot directly observe judgement, reasoning or critical thinking. They observe presentations, reports, recommendations, analyses and decisions. Those outputs became the evidence through which organisations assessed performance.
For most of the history of knowledge work, this approach made sense. Producing good work requires thinking. Researching, analysing, writing and deciding were inseparable activities performed by the same person. The quality of the output reflected the quality of the judgement behind it closely enough for organisations to use one as a proxy for the other.
Artificial intelligence changes that relationship.
Today, two employees can produce work of almost identical quality while contributing very different levels of thinking. One may use AI to generate an initial draft before questioning assumptions, testing alternatives and refining the final recommendation. Another may accept the first plausible answer and submit it with minimal reflection. Their outputs may look equally polished, yet the capability demonstrated by each employee is fundamentally different.
AI adoption is no longer the central challenge. The harder problem is measuring human capability when AI increasingly contributes to the work. Organisations will need to rethink one of the assumptions that has quietly underpinned performance management for decades: that output reliably reflects capability.
Why Output Became the Default Measure
Performance management evolved around what managers could reasonably observe.
Peter Drucker argued that the central challenge of the knowledge economy was improving the productivity of knowledge workers. Unlike manual work, knowledge work produces intangible value through thinking, analysis and decision-making. Managers cannot watch someone think in the same way they can observe a manufacturing process. Instead, they evaluate the outcomes of that thinking through the work people produce.
Over time, organisations built systems that rewarded visible contribution. Performance reviews assessed deliverables. Promotions recognised consistent results. High-potential employees were identified by the quality and reliability of their work. While no manager believed output captured everything, it was generally a reliable indicator of capability because producing excellent work required excellent thinking.
The assumption was imperfect, but it worked for many decades.
AI did not suddenly make outputs an imperfect measure of capability. They always were. What AI has done is expose the limitations of relying on outputs alone, making the gap between visible work and human judgement impossible to ignore.
The arrival of AI breaks that relationship, forcing us to rethink how we measure performance and potential.
When Thinking and Output Separate
As AI becomes capable of producing increasingly sophisticated work, the relationship between human effort and visible output begins to change.
This shift has important implications for how organisations evaluate performance.
If AI can produce competent work quickly, the finished output tells us less about a person’s contribution than it once did. A well-written strategy document may reflect deep analysis, careful judgement and iterative collaboration with AI, or it may rely on far less independent reasoning and oversight. Looking only at the final document makes those situations difficult to distinguish.
This does not diminish the importance of results. Organisations still exist to create value, and customers ultimately experience products and services rather than thought processes.
The Shift Towards Decision Quality
Historically, outputs were a reasonable proxy for capability because producing high-quality work required high-quality judgement. AI has changed that. Today, the same polished output can reflect deep expertise, effective use of AI, or very little independent judgement. Outputs still matter, but they no longer tell the whole story.
This is where decision science offers a useful perspective. The distinction between a good outcome and a good decision is not new. For decades, strategy consulting, behavioural economics and military leadership have recognised that outcomes alone are an unreliable way to judge the quality of thinking.
Daniel Kahneman described this tendency as outcome bias: our inclination to judge a decision by how it turned out rather than by the quality of the reasoning behind it. A well-reasoned decision can produce a poor outcome because of factors outside our control, while a flawed decision can occasionally produce a successful result through luck.
AI makes this distinction impossible to ignore. When polished outputs can be produced with varying levels of human judgement, organisations can no longer infer capability from outcomes alone.
Managers should not simply ask whether an employee produced a good report or whether a recommendation happened to succeed. They should also ask whether the employee exercised sound judgement throughout the process.
Did they challenge assumptions?
Did they recognise uncertainty?
Did they verify important information?
Did they know when AI was helpful and when human expertise was essential?
These questions reveal capabilities that the final output alone cannot.
What Leaders Should Measure Instead
This does not require organisations to replace performance management systems overnight. Instead, it requires broadening the conversations managers already have.
When reviewing AI-assisted work, leaders should become more interested in reasoning than process. Asking someone how they approached a problem often reveals far more than asking how quickly they completed it. Understanding why they accepted one recommendation and rejected another provides insight into judgement that cannot be seen in the finished document.
Managers should also pay attention to calibration. Strong performers know when AI provides sufficient confidence to move forward and when additional evidence, expertise or discussion is needed. Neither blind trust nor complete distrust of AI represents effective judgement. The ability to calibrate confidence may become one of the defining professional skills of the coming decade.
Another important capability is explainability. Employees should be able to explain not only what they decided but why they decided it. If reasoning cannot be articulated, it becomes difficult to evaluate, improve or transfer to others. As AI becomes more deeply embedded in workflows, transparent reasoning becomes increasingly valuable for organisational learning.
Finally, leaders should reward intellectual curiosity. The strongest AI users are rarely those who generate answers most quickly. They are often those who ask better questions, explore alternatives and remain willing to challenge convincing but incomplete responses. These behaviours strengthen both human judgement and organisational capability over time.
The Organisational Challenge
The implications extend well beyond individual performance reviews.
Promotion decisions, succession planning, leadership development and hiring a talent all depend on recognising capability before it becomes visible through senior responsibility. If organisations continue to rely heavily on outputs that AI increasingly helps create, they may struggle to distinguish future leaders from employees who simply use technology effectively.
This is ultimately an organisational performance issue rather than a technology issue.
Organisations are likely to realise the greatest value from AI when they redesign work, not simply automate more tasks. As AI takes on more routine execution, people can focus on activities where human judgement, critical thinking, creativity, ethical decision-making and relationships add the most value. These capabilities have always been important, but as AI assumes more of the execution, they become easier to identify, develop and evaluate.
Performance management therefore becomes less about measuring activity and more about understanding how value is created.
A Question Worth Paying Attention To
One question continues to stand out, and I don’t think we have a complete answer yet.
If judgement becomes one of the most valuable capabilities in an AI-enabled workplace, how do organisations intentionally develop it?
For generations, judgement developed through experience. Early-career professionals wrote first drafts, made imperfect decisions, received feedback and gradually learned to navigate uncertainty. Those routine tasks were more than productive work. They formed the apprenticeship through which judgement, professional taste and decision-making were built.
As AI takes on more of that work, organisations risk removing the very experiences that developed those capabilities. The issue is not simply that junior roles are changing. It is that the pathway through which people learn judgement is changing with them.
This demands a different approach to both development and performance evaluation.
Not every task that AI can automate should be automated. Some work creates more value as practice than as production. Early-career tasks often exist not because organisations need another draft or spreadsheet, but because they develop judgement, professional taste and decision-making. Automating them entirely risks removing the apprenticeship that future leaders depend on.
At the same time, organisations may need to separate the evaluation of AI from the evaluation of employee performance. Rather than measuring people primarily by the quality or speed of the final output, leaders should place greater emphasis on the quality of reasoning behind it.
That means asking different questions.
Why was AI used in this way?
Which assumptions were challenged?
What information was verified?
Where did human judgement override AI?
What limitations or errors were identified?
How did the employee respond when AI produced an incomplete or misleading answer?
Those conversations reveal capabilities that the final output alone cannot. They provide a far stronger signal of future leadership potential than whether a report happened to be well written.
Organisations will adapt to AI more successfully when they develop human judgement alongside technology and recognise it through the way performance is measured and rewarded.
The challenge is no longer whether AI can produce increasingly capable work. It is whether organisations can continue developing people who know when AI is right, when it is wrong, and when it should not be trusted.
That may become one of the defining leadership challenges of the AI era.
If you found this article useful, please consider sharing it with a colleague, manager or leader who is thinking about how AI is changing the way we measure performance.
Thank you for reading.
*This essay is intended to be read alongside Charles Kent’s companion article in Agent Autopsies, which explores how AI is changing the execution of knowledge work. Together, the two essays examine how AI is reshaping both the work itself and the way organisations evaluate human capability.
Clarice x
Clarice is currently building The Human-First Shift, a people intelligence platform designed to help organisations understand and predict performance in the age of AI. The platform helps leaders identify people-related risks and estimate their financial exposure to hidden organisational costs such as burnout, disengagement, poor leadership, attrition, low productivity and ineffective AI adoption.
Clarice De Chavez is a Talent Strategist helping organisations navigate workforce transformation in the AI era. As Founder of The Talent Seed and Co-founder of The Human-First Shift, she advises on future-ready talent, leadership, and hiring. With nearly 20 years’ experience across global brands and European fintech scale-ups, she writes about AI, hiring, leadership, career transitions, and the future of work.
You can find Clarice on LinkedIn.
Discover Clarice’s Essays
Everyone Is Using AI at Work. Almost No One Is Being Trained.
The Middle Manager as Translation Layer: What the Data Actually Shows About AI Adoption
Why Your Job Search Is Stalled (And It’s Not a Skills Problem)




Love it! I really hope both articles together serve to start the conversation about performance management in the AI era!
Great essay, Clarice! Very deeply thought and very sharp, pointed and practical suggestions. Love it!