AI and Technology
Blog: What happens to job design when we can test a role before we build it?
Author: Ali Akkaya, Senior Digital Solutions Manager, CRF
In a recent business strategy meeting, my colleague Karen – the MD of our sister company, PARC – said something I haven’t been able to put down. She said she doesn’t ask AI how to do her job. She asks which parts of her job should be AI’s at all.
Read that back. A business leader, not a technologist, had just described which might be the most important skill of the next decade, and she’d have called it common sense.
Because for sixty years, the science of work has quietly known how to answer Karen’s question. Decades of research, from the Tavistock coal-mine studies onwards, converged on one stubborn finding: performance is a property of how the work is designed, not of the worker, and not of the tool. The theory was settled. The experiment was the problem. Redesigning a human role takes months and lands on real people, so mostly we guessed, restructured, and read the post-mortem three years later.
Last time, I asked whether we’re building AI agents the same careless way we’ve built jobs for sixty years. This piece is the other half of that thought. Because GenAI has just handed the science of work the one thing it never had: a way to test a role (with scope, tools, authority, hand-offs, checks) before we commit a single human being to it. Stand it up, tear it down in minutes, watch exactly how it performs. Aviation had a machine (wind tunnels) that did this for flight, and I’ve leaned on that image before; borrow it if it helps. But the machine matters far less than what it lets us finally do: design the work on purpose.
So, here’s the flag I want to plant, and I’ll be blunt about it. The biggest risk in this transition isn’t the technology. It’s who we’ve left holding the design. If it’s IT leading the automation in your company, you might want to think again. Automation is an operating-model transformation, not an IT project. IT can lead the tooling, but designing the work is a business and people decision, and the leaders best placed to make it are too often the ones still waiting to be told what the tool can do. Technology is the enabler here. It was never meant to be the author.
A forty-year-old prediction, coming true at scale
Before anyone tells you GenAI’s effect on work is unprecedented, let me introduce Lisanne Bainbridge. In 1983 she wrote a short paper, Ironies of Automation, and its central irony has aged frighteningly well. Automate the hard parts, the thinking went, and people are left with the easy ones. What actually happens is the opposite. Automation strips out the easy, routine work and leaves the hard residue behind, more abstract, more vigilant, heavier per decision.
Four decades on, BCG put numbers to her hunch. Two-thirds of regular AI users, 67%, say their job satisfaction went up. And 41% say their cognitive load went up too. Wait a minute, you wouldn’t normally expect those two to move together, would you? Higher satisfaction usually comes with less strain, not more. Both are true, though, and once you think in job-design terms you can see why. AI clears out the repetitive parts of a role and leaves the residue: the ambiguous, judgment-heavy, someone-has-to-decide work. And let’s not forget the chore nobody put in the budget: checking the machine’s homework which is why nearly half of users now spend more time directing and reviewing AI than doing the task themselves.
In Hackman and Oldham’s language, the very same change can enrich a job (more skill variety, more autonomy) or hollow it out (less sense of the whole task, relentless load). And which one you get is not a property of the model. It’s a property of the redesign you did, or didn’t do.
So, these tensions aren’t contradictions to argue about. They’re design variables. And here’s where I get a bit stubborn: those variables belong on a business leader’s desk, or a job designer’s, not the IT department’s (alone), and not the desk of whoever happens to be building the agents. This is HR’s fight, and the C-suite’s. If the people who understand the work aren’t in the room when the work is redesigned, don’t be surprised when the redesign forgets the human in it.
Which brings me back to Karen’s point: If most of your organisation is still asking “how do I write a better prompt?”, they are, in effect, asking the machine how to do their jobs. And I get it, how you prompt a tool does matter and there is definitely value in it. But what matters far more is the question you ask yourself, as a business leader, long before you ever open the tool.
Don’t ask AI how to do your job. Ask which parts of the job should be AI’s at all.
That one reframe does most of the work. It turns a technology question into a design question, and it puts you, not the model, back in charge of the answer. Map the process first, then interrogate it honestly. What are you actually trying to achieve, and would you or your clients still want it if the tool didn’t exist? Which parts of the work are routine and rule-bound enough to hand over? And which parts turn on judgment, carry the weight of accountability, and should stay firmly with a person? Hand over the first kind. Guard the second. That is the job, not prompt-craft.
The gauges lie, unless you fit real ones
Here’s the catch. We built the tunnel before we learned to read the instruments. In 2025, METR ran a proper randomised controlled trial, experienced developers, real tasks, codebases they knew well. Going in, they expected AI tools to speed them up by 24%. In reality, they were 19% slower. And afterwards, they still believed they’d been roughly 20% faster. That’s a forty-point gap between what people felt and what actually happened which should worry anyone whose AI business case rests on asking staff how much time they think they saved. Which, let’s be honest, is most AI business cases.
And the failure isn’t loud, either. BCG’s field experiment with 758 consultants mapped what they called the “jagged frontier”: just outside the boundary of what AI does well, on tasks that looked almost identical, consultants using AI were 19 percentage points more likely to be confidently wrong. AI doesn’t fail like a machine. It fails like a plausible colleague.
Think about your own experience with GenAI for a second. From time to time, hasn’t it felt like that? Sometimes the answer reads beautifully and sounds completely plausible, and only when you stop and look harder you notice it’s wrong, or simply not adding any value.
Autonomous does not mean unsupervised
If I had to defend one line of the AI budget hardest, it would be instrumentation: the ability to see what the work is actually doing, continuously, from the inside. Taylor measured work once, from the outside, with a stopwatch. A traced agentic workflow measures it all the time, and the same signal tells the engineer whether the agent is drifting, the risk officer whether decisions can be audited, and the executive whether the freed-up capacity actually went where it was promised. The practitioners are already there: in LangChain’s survey of 1,300+ agent builders, 89% had this kind of observability in place. Keep one simple rule in mind: a goal you can’t measure is just a hope, and measuring with no goal in mind is just a dashboard nobody acts on. And you need both.
So, what will you do with the freed time?
The time savings, at the level of individual tasks, are real. BCG’s 2026 survey of nearly twelve thousand workers found that 42% of regular frontline AI users save a full working day or more, every week. A working day. Per person. Per week. And then what? Nothing. Two-thirds of them said they get little or no guidance on what to do with that time, and more than half don’t redirect it to anything of higher value.
So where does it go? Cyril Northcote Parkinson answered this back in 1955: work expands to fill the time available. GenAI has just shipped the upgraded edition: work expands to fill the capacity you’ve freed. An eight-month Berkeley study watched it happen in slow motion: expectations quietly rise, scope widens, and the extra fatigue never shows up on a dashboard. The freed day doesn’t leak away. It gets absorbed, silently, into a slightly heavier version of the same job, unless someone decides otherwise.
Because freed capacity always has a destination, and if you don’t choose it, Parkinson chooses for you. The usual menu – more output, higher quality, innovation slack, skill development, wellbeing, headcount – is shorter than it should be. I’d add four that rarely make it onto the slide: adjacent revenue, price, reskilling, and resilience. Each of those deserves a piece of its own, and I won’t pretend I’m the expert on all of them. But here’s the one thing I’m sure of: a destination without redesign dissipates. Same law as everywhere else in this piece.
Two Swedish companies make the whole argument better than any survey could. Klarna announced in early 2024 that its AI assistant was doing the work of 700 human agents. By May 2025, its CEO admitted the cost-cutting had “gone too far” and started hiring humans back; the quality that underwrites a brand had quietly walked out with the headcount.
IKEA ran the experiment properly. Its chatbot, Billie, resolved around 47% of customer enquiries, respectable enough. But the decisive move was what Ingka (IKEA’s parent company) did with the 53% it couldn’t handle. They mined those unresolved queries and found them clustering around interior design – customers essentially asking, will this sofa actually work in my living room? That pattern became the business case to reskill 8,500 call-centre co-workers into remote interior-design advisers, and the new channel generated €1.3bn in its first full year. That, in a single move, is the thesis of this piece: the chatbot’s failures were the experimental data. Ingka didn’t ask AI how to do the job. They read what it couldn’t do as a map of where human judgment still carried value, and redesigned the roles around it.
My honest caveat, and I try to keep one in every piece, because the discipline I’m preaching has to start at home: if I can’t say where my own argument breaks, I haven’t finished thinking. IKEA is not a fairy tale. It made significant office-role layoffs in 2026, unrelated to this on its own account, which is a whole other discussion happening right now, and a reminder that reinvestment is a decision, not a permanent moral guarantee. Klarna’s “700 agents” was workload-equivalence during a growth phase, not 700 redundancies. Neither case hands you a template. What they hand you is thequestion, and the evidence that answering it deliberately, in advance, is what separated the two outcomes. MIT found that roughly 95% of enterprise GenAI initiatives return nothing, while the redesigned minority captured almost all the value. If you know me a little bit you’d know the hill I’ve chosen to die on: start with the why, not the tool. The differentiator was never the model – I mean, it can very well be, but that’s not the point here- it was whether anyone had decided, before the licences landed, what the work was actually for.
So, where to start?
None of this is complicated, but almost none of it happens on its own. Each of these could carry a piece of its own (and some will) so for now, just the headlines:
- Start with the why, not the tool. Decide which workflow you’re changing, where the freed time goes, and the one number that proves it worked. No destination, no deployment.
- Redesign the process, not the task. Map where the work actually waits and gets stuck. A pile of licences is not a strategy.
- Measure value, not adoption. Hours redirected, not hours saved; they are not the same number.
- Aim at the jagged frontier and re-check quarterly. What AI does well keeps moving; so should your split of work between human and machine.
- Make oversight mean something. Watch one number: how often a human overrides the machine. It’s an honest question about the job you’ve designed: is it a real job, or something darker?
On that last question, let me be plain. The unhealthiest job design occupational psychology knows of is the one where everything depends on you and nothing is up to you. That is the default job description of the human in the loop, unless leaders design it otherwise. There’s a name for it, it comes from the anthropology of automated-vehicle crashes, and it’s the subject of the next piece: the moral crumple zone.
CRF’s AI in HR Series
This blog is part of the CRF AI in HR Series which is built on a simple premise: the hardest part of AI in HR isn’t the technology, it’s knowing where to use it and how to create value.
HR’s role is not to become the AI expert, but to help the organisation make better decisions about work, people and capability.
Across four parts, the series focuses on creating value with AI, enhancing individual productivity and judgement, enabling line manager effectiveness and redesigning work and capability. It brings together live webinars, on-demand learning, cohort-based programmes, hackathons, practical research and articles to help HR teams move from AI curiosity to AI capability.
