Why personal AI practices don't scale to the team

Sergey Golubev 2026-07-23 24 min read
🌐 Читать на русском

Illustration: an engineer holds a productivity lever down at the x1 notch, though it could go higher. Above him looms a teetering avalanche of tickets about to fall. Personal AI practices barely scale to a team. One person speeds up, while the team around them stays at the old pace. At my own company we’ve gotten a bit further, but we’re still in the middle of that adaptation. I’ve collected the observations in one place: why this happens and what the people who make it work actually do.

The main reasons teams stall

Team resistance: the root of the problem

Rolling out AI in a group almost always runs into resistance. The reasons repeat: fear of being replaced, skepticism from past bad experiences with tools, being overloaded with current work. People are afraid AI will make their skills redundant or, worse, expose that they aren’t as good as they seemed. In companies over 300 people this is most pronounced.

The scale shows up in the numbers. In a Writer and Workplace Intelligence survey (1,600 people, 800 of them executives), 31% of employees openly admitted they sabotage their company’s AI strategy: refusing to use the tools, deliberately producing weak output, gaming the metrics. This isn’t passive “haven’t figured it out yet,” it’s active pushback. In the updated survey in spring 2026 the number barely moved.

But the word “resistance” often hides something other than what it looks like. A common story many team leads describe: the loudest skeptic on the team turns out to be an active user of the tools. Public skepticism works as a negotiating stance, not as a technical assessment.

And the grounds for that skepticism are usually rational.

The first is about quality. The share of code the author can’t explain keeps growing. Not “bad code,” but genuinely unexplainable: it works, the tests pass, and when you ask “what’s going on here” you get a shrug. What experienced engineers object to isn’t the use of AI, but that the habit of thinking about the result disappears along with the speedup.

The second reason is about workload, and I’ll come back to it below.

There’s a separate problem people mention less often: review turns into the bottleneck, and the load on whoever owns it grows faster than anyone else’s. Generating code gets cheaper; checking code doesn’t get cheaper at all. Team leads and senior engineers hit the ceiling first, because everything the team managed to generate passes through them. The skepticism of the person hand-reading that stream is data, not sabotage.

One more thing worth admitting honestly: training almost always happens on the employee’s personal time. Evening calls, weekends, “take a look when you get a chance.” If you don’t say this out loud, it becomes quiet coercion that later comes back as resistance.

Uneven readiness and the difference in depth

Every team has a different level of “nativeness” to new tools. Some are ready to experiment and don’t give up after failures; others resist. Trying to lift everyone at once scatters the effort and produces nothing.

Readiness diverges not only across people but across functions, and not always predictably. Often the first to catch on aren’t the developers but the testers or analysts. Their trigger is usually economic rather than technological: the function has fewer people, the workload has grown, and automating the routine becomes a matter of survival, not curiosity.

Another observation, this one from personal experience and tested on more than one person: handing out links doesn’t work. I’ve shared carefully curated lists with people plenty of times - channels, videos, courses where you can genuinely learn. Those collections get consumed like a feed and forgotten. Watched it, got a dose of new information, felt better.

It’s not about the quality of the material. People simply don’t know how to work with it systematically, and they resist a flow they aren’t ready for. Someone glances into the collection, realizes they aren’t ready to dive in headfirst yet, and leaves it at that. Between “learned about it” and “applied it” lies work no one will do for them, and a single link doesn’t kick it off.

Now about loudness. The first and most visible “experts” on AI in a team aren’t the ones who understood it more deeply. The people who actually grasp the limitations usually take longer to figure things out and speak more quietly. That lowers the quality of internal expertise and undermines trust in the topic itself: the team sees confident claims, then sees their consequences, and draws a conclusion about the technology rather than about the person speaking.

And still, the loud ones have a point, and it’s worth acknowledging. It used to be that to get the tool or feature you needed, you had to go through a long official path: write up requirements, create a ticket, drop it into the backlog, and wait until the developers got to it. That takes weeks, and often the task just stays at the bottom of the list. The tools broke that queue: a person just builds the prototype themselves, in an evening, without waiting on anyone.

And here the loud ones are right on the merits. If someone built what the official channel would never have delivered, their initiative is justified - even if they picked the wrong tool and it will have to be redone later. The complaint “they chose the wrong tool” is fair, but it doesn’t cancel the main point: the old backlog path simply didn’t work for them.

The problem of going from personal to shared

The prisoner’s dilemma of AI adoption

An individual engineer often has no incentive to reveal their full potential. By becoming an x10 engineer, they risk getting x10 the workload and ultimately helping to cut colleagues. A collectively destructive equilibrium forms: everyone rationally understates the productivity they show.

The logic is simple and takes a minute to work out. Show double the speed - get a doubled plan next quarter. Take initiative - it gets picked up, carried upstairs, and you’re the one accountable for it. Stay quiet - keep what you have, and pocket the gains from the tools as free time.

This isn’t one team’s observation. Anthropic used its research tool to interview 1,250 professionals: 69% mentioned in some form the stigma around using AI at work, and 55% are anxious about what happens to their careers next. People hide it not because they’re embarrassed by the technology. They’re reading what admitting it would mean.

The backdrop for that calculation is very real too. In February 2026 Block announced it was cutting more than 4,000 of roughly ten thousand employees, naming as the reason directly that a smaller team with new tools does more. The stock rose more than twenty percent. An engineer reading that news draws exactly the conclusion described above.

Hence a non-obvious practical takeaway: frame the training as the person’s own development, not as the company’s investment in productivity. The framing matters more than it seems. The moment it becomes “the company invested, now pay us back in speed,” you trigger precisely the dilemma that makes people stop showing their real results.

The individual gap and gamed metrics

The main unsolved problem is the gap between the individual benefit from AI and the lack of benefit to the organization. The developer speeds up; the business metrics don’t move. Paul Everitt, in his talk on the shift to agentic engineering, puts it plainly: we see the individual benefit, the organizational one not yet.

He also has a number that sobers you up. DX observed 400+ engineering organizations for 16 months: use of AI tools grew by an average of 65%, while the median gain in pull-request throughput came to about 8%. Everitt’s phrasing: it isn’t ten times, it’s about ten percent. A follow-up to the same study has something even worse: even these modest gains are unstable, and two-thirds of the developers who reached peak time savings slide back in the following quarters.

Macroeconomists tell a similar story. Daron Acemoglu, in his paper on the simple macroeconomics of AI, gives an upper-bound estimate for the gain in total factor productivity of less than 0.7% over ten years. Not per year.

At the organizational level, the 2026 picture is the same. In the January edition of Accenture’s Pulse of Change (3,650 executives and 3,350 employees, 20 countries), only 32% of leaders say they’ve achieved a sustained, enterprise-wide effect from AI. From the employees’ side it’s symmetrical: only 27% firmly confirm they’re comfortable delegating tasks to AI agents. The investment flows in, and the sustained effect belongs to a third.

I’ll cite one more slice with a caveat, because it’s a survey by an AI-platform vendor, not independent research. Writer, together with Workplace Intelligence, surveyed 1,200 executives and 1,200 employees in April. 79% of organizations face problems with adoption, and only 29% see meaningful returns from generative AI. Meanwhile 92% of leaders are deliberately cultivating an internal “AI elite,” and 60% plan to lay off those who don’t adapt. One note: the claimed fivefold productivity of that “elite” isn’t a measurement, it’s the opinion of 87% of the surveyed leaders.

And here two blocks of numbers snap into one unpleasant picture. There’s almost no return at the company level, yet layoffs for failing to adapt are already being planned. The engineer gets confirmation of their calculation from both sides: showing a productivity gain is risky, and not showing one is risky too. This isn’t irrational fear, it’s reading the room.

So why doesn’t personal acceleration turn into organizational acceleration? Because the speed of the whole chain is set not by its fastest stage but by its slowest. By speeding up the writing of code, you accelerate one segment, but the overall timeline is set by the one that stayed slow. And as soon as you clear one jam, the queue immediately forms at the next stage - in planning, in review, in testing. The place that throttles the whole flow doesn’t disappear, it moves. And it moves faster than anyone notices: the team is delighted that coding got twice as fast, while releases somehow keep coming out at the old rate.

First: a significant part of the gain goes into preparation. The more you lean on the agent, the more time goes into framing the task, gathering context, and reviewing the plan. The work doesn’t vanish, it changes shape, and in the reports it looks like “coding got faster, but the release somehow didn’t.”

Second: checking doesn’t scale along with generation. You can write code many times faster; testing and reviewing it, hardly at all. If you don’t rebuild the checking function in advance, faster development simply creates a queue in front of QA and burns people out there.

Now about metrics, and this is the most absurd part of the story.

The classic anti-pattern looks like this. Inside the team everyone understands that story points have long lost their meaning: with agents, a three-point and an eight-point task get done in about the same time. But upstairs they need to see growth, so the metric gets engineered. For example, estimating new tasks against the old complexity guide - then velocity is guaranteed to rise, because yesterday’s two-week task now closes in a couple of days at the same point cost.

Then this logic reaches its limit: the estimating gets handed to the model itself. You end up with a closed loop where AI generates the metric that’s supposed to prove the effect of adopting AI. The hole is obvious right away: growth is measured in units that, in the new reality, no longer measure anything.

This is also where you should handle the well-known MIT figure carefully - 95% of GenAI pilots with no measurable impact on P&L. The number is real, the report is real. But most pilots simply had no measurement before rollout. “No measurable effect” in that situation is the default outcome, not a proven failure. Which brings us back to the same point: first agree on what you measure and how, or any number you get will say more about your accounting than about your rollout.

If you measure individuals, behavior gets distorted: some draw activity, some slow changes down so as not to spoil their personal numbers. A good replacement shows up in different teams: a leaderboard by tokens spent didn’t work, and it was replaced with a leaderboard of process improvements - points for building a shared tool, for reporting a bug in the shared rules and agents. Measure activity and you get performed activity. Measure contribution to the shared tool and you get a shared tool.

And an important correction to the thesis itself. Metrics get distorted not only from below and not only out of self-interest. Often the number is carefully engineered by the team’s own management - as a shield against pressure from above, to buy time and not expose the people. That changes the recipe: what you have to fight isn’t the people fudging the numbers, it’s the demand to present growth where there’s nothing yet to measure it with.

Lack of sync and standards

When everyone uses their own AI with their own settings and context, the results diverge: different tone, different assumptions, different quality. It gets worse when those results feed each other’s work. Until a team has a shared layer for the agents - common rules and instructions, reference docs, guardrails, reusable skills, described processes, and agreements on how everyone works with the model - AI amplifies the inconsistency rather than dampening it. Each person tunes the agent to themselves, and the output is five different styles instead of one team style.

The essence of the problem fits into one question: if five developers don’t share the same context, how big will the spread be in quality and in the way the result is formatted?

The first solution that comes to mind usually doesn’t work. Setting up a shared repo where everyone dumps their own work isn’t standardization. Pretty quickly you get a junk pile where each person uploads what worked for them personally, and making sense of it is harder than writing it from scratch. The same inconsistency, just centralized.

What works better is a described shared context: architecture, conventions, the data model, rules for working with external dependencies, anti-patterns, security requirements. Separately, context for testing: test-case templates, bug-report templates, automated-testing standards. Plus a designated owner who keeps that context current and reviews changes to the shared tools. It’s better to make the role time-limited, say for a quarter, and you must give the person somewhere to log that time. A role on top of a full workload doesn’t survive.

One more effect shows up only when a team starts taking apart each other’s tools: people duplicate shared reference docs inside their personal work. The list of roles, the list of statuses, the naming rules live inside one tool, and when a new entity gets added tomorrow, you have to edit it in several places at once. Hence a useful split into infrastructure things, which one person builds and maintains while everyone plugs them in, and personal things, where each person is free to do as they like.

And an honest caveat against myself. Standardization isn’t an unqualified good, and resistance to it is sometimes justified. The argument “give everyone the freedom to build for themselves and let’s see what comes of it” makes sense at an early stage, when it’s still unclear which practices are even worth locking in. The pattern here is this: a standard lands when it’s proposed by a technically competent person who uses it themselves. The same standard handed down administratively reads as control, and the resistance will be political, not technical.

The skills gap and shifting requirements

Resistance to traditional practices

The move from solo work to a team is a critical moment. The instinct to impose familiar order through feature branches, code review, and scrum sharply reduces speed. This is the main problem of team AI coding.

Then come decisions that not long ago would have looked like process regression.

Sprints start getting in the way. The team thinks in a two-week horizon with a fixed set of tasks to ship, and it loses any reason to close a task in a day. Detailed grooming also loses its point: manually breaking down what’s going to be decomposed with the model anyway is a waste of time.

The most counterintuitive thing happens with testing. Splitting a feature into small testable pieces is canonical practice, but with fast generation it produces double the work: you test it in parts, and then re-check the assembled whole anyway. Many teams end up rolling back to waiting for finished functionality and testing it once. Formally this is a step back to waterfall. And it starts with testing, because that’s exactly where accelerated generation piles up, but development follows: planning grows coarser, tasks stop being cut into small iterations.

And one more finding people don’t reach right away: WIP limits at the checking stages work as a defense against pressure. When the input got faster but the output physically can’t keep up, the limit gives the business a clear line to hear. A process artifact turns into an argument that keeps the speedup from burning out the people at the end of the conveyor.

New requirements for team members

Purely technical skills are getting cheaper relative to understanding the business domain. Engineers need business sense and empathy for the user. The “my job is to write code” mindset becomes a brake.

It’s most visible in how task-framing changes. It used to be that broken-down user stories reached the developer, meaning someone else did the main decomposition work. Now, more and more often, what comes in is a PRD - a document not taken down to the level of individual tasks, and the engineer does the detailed planning themselves together with the model. For them it’s noticeably harder: more questions, more uncertainty, everything needs to be investigated and clarified. But the business context ends up in the head of the person writing the solution, instead of being relayed secondhand.

But here it’s worth challenging a popular thesis. The idea that engineers themselves hide from talking to the business is far from always true. Time and again it’s the opposite: technical specialists are systematically not invited to meetings with the client or to discussions where decisions get made, and to get their time you have to negotiate through several layers. The barrier between engineer and business is more often organizational than mental. Before fixing people’s mindset, it’s worth checking whether they’re even let into the room.

Ecosystem barriers and practical norms

Tacit knowledge and blocking by IT

The main bottleneck for adoption in large organizations is tacit knowledge, which is extremely hard to extract. It seems obvious to the people who hold it, so it doesn’t get spoken or documented. Against a backdrop of forced rollout and talk of layoffs, active resistance appears that blocks the process further.

There’s a fear here that sounds paranoid but has grounds: companies will digitize their own people. They’ll give people an interface while keeping the inner workings to themselves. People work with the tool, training an internal agent as they go, and one day they turn out to be surplus.

The reality, though, looks duller and is more interesting for it. More often it’s not malice but the ordinary organization of work: the client assembles all the scaffolding on their side and gives contractors access to the finished thing. The contractors do the work but don’t see how the tool is built and can’t change it. No one built a trap. Ownership of the context simply stays, naturally, with whoever assembled it, not with whoever uses it. This is structural asymmetry, not a conspiracy, and you have to defend against it differently.

The flip side is worth acknowledging too: extracting tacit knowledge is hard even with a full willingness to share. Often there’s simply nothing to share yet, because the experience hasn’t settled into a state you can put into words. Add the language barrier in distributed teams, and half the context doesn’t reach the people who need it.

The problem of personal AI tools

IT departments often block personal AI tools on security grounds without realizing the cost. An employee with a personally tuned AI works more effectively than with a clean corporate deployment of the same model. The reason is simple: the personal setup accumulates context around a specific person and their tasks, while the corporate one is handed to everyone the same.

It’s important not to overstate this. I’ve come across an estimate that blocking personal tools costs a company a multiple loss of productivity, but I couldn’t find a single study behind that figure, so I won’t cite it. What has actually been measured: in a Gartner survey (12,000 respondents, 40 countries), employees who combine the corporate and personal stack are 1.7 times more likely to report meaningful time savings than those who use only the approved corporate tools.

And the main thing: a ban doesn’t remove the use of personal tools. It removes its visibility. When corporate limits run out mid-day, people switch to personal subscriptions and keep working. The reaction to this is usually not “let’s revisit the limits” but “don’t tell the others about this.” The problem doesn’t go anywhere, it just drops out of sight of the people who could solve it.

Add to this the difference in terms. A corporate plan usually spells out what the provider can and can’t do with work data. A personal subscription has its own rules, often less strict. Figuring out where the data ultimately goes and how safe that is falls, in practice, on the employee themselves. And they have neither the time nor the authority for that analysis - it’s the security team’s job, not the developer’s.

One more reason people forget: often it’s not about psychology at all, but about plain infrastructure. A person can’t create an account because the service won’t accept their country’s phone number. Access is closed by region. The corporate network cuts off the necessary traffic. There aren’t enough licenses for everyone who wants one. Until these basic things are solved, the conversation about motivation and fears simply never starts - the person hits a wall at the entrance.

Hence a clear market niche: a product that can separate the professional and personal context while staying inside the security perimeter.

How to work with this

Now for what actually delivers results.

Working directly with fears. If you don’t address, head-on, the fear of replacement, the fear of exposing incompetence, and the overload, the rollout will fail. This isn’t a soft add-on to the process, it’s the price of entry.

A same-level cohort instead of “lift everyone.” A small group goes through the training together and comes out with a shared understanding and a shared vocabulary. From there it grows as a group and becomes a source of practices for the rest. Trying to cover everyone at once produces the exact opposite effect: the strong are bored, the weak are lost, and it takes more time.

The “you first, company second” frame. First the person’s own development, and only second the benefit to the product. This framing removes both the fear and the prisoner’s dilemma better than any persuasion, because it honestly answers the question “what’s in it for me.”

A regular sync instead of a one-off event. Once a week the team gets together and swaps experience: who tried what, what worked, what didn’t, which findings and skills they’re ready to share. The exchange format matters more than the lecture format. One-off training doesn’t work this way: it leaves people no time to let the knowledge settle and try things themselves, and without your own attempt nothing sticks.

A public review of someone else’s work. One person shows their tool on a real task, and the team takes apart how it’s built. This gives you three things at once: the work gets better, the others see it from the inside, and the shared pieces get moved to where everyone can use them.

And the most interesting part about resistance. Skeptics often turn out to be the best critics in these reviews: they’re the most attentive and ask the questions the enthusiasts skip. Resistance isn’t converted by persuasion or by showing off someone else’s success. It’s converted when you give the skeptic the role of expert in someone else’s work, rather than the role of a target to be convinced.

A measurable everyday win as the argument. Not abstract productivity, but a concrete routine that used to take hours and now takes minutes. That kind of example convinces more than any presentation, because you can repeat it tomorrow.

Promising the business “quality, not x10.” What you sell upstairs shouldn’t be a growth in volume but the freed-up time for refactoring, tech debt, and automating checks. It’s the one framing that doesn’t trigger the prisoner’s dilemma.

A focus on the team result. Instead of individual speed - the shared result: time-to-market, the share of releases that get rolled back, the number of bugs after release. And shared rules as the way to standardize.

Separately, on what not to do. I once served as a mentor at an AI hackathon at a large company: strong instructors and mentors, the company invested seriously. And still it didn’t take hold as a one-off - for one reason. People weren’t freed from their main work. Some teams simply never made it to the hackathon because of their workload, and the projects that did start weren’t finished, and everyone went back to the routine. They gave money, they didn’t give time. That’s the main takeaway: if you aren’t ready to take part of a person’s current load off them, any rollout - a hackathon, a course, a sync - stays their personal evening hobby.

The bottom line

Personal practices don’t become shared not because people are lazy or haven’t figured it out. They don’t become shared because the person has no incentive to give them away, the team has nothing to measure the effect with, and the organization asks for growth where the bottleneck has already moved somewhere else.

First take the routine and the fear off people. Then agree on what exactly you’re measuring. And only then deal with the tools.

Sources

  1. Paul Everitt: The Shift to Agentic Engineering - the gap between the individual benefit from AI and the missing organizational effect.
  2. AI productivity gains: More modest than expected - DX, 16 months and 400+ engineering organizations: a gain of about 8-11%, not multiples.
  3. The AI efficiency plateau - DX on why individual gains slide back.
  4. Introducing Anthropic Interviewer - 1,250 interviews: stigma around using AI at work and anxiety about careers.
  5. 31% of employees are ‘sabotaging’ your gen AI strategy - the forms active resistance to a rollout takes.
  6. Enterprise AI adoption in 2026 - Writer and Workplace Intelligence, April 2026: adoption problems, ROI, the “AI elite.” A vendor survey, read with that in mind.
  7. Accenture Pulse of Change - January 15, 2026 edition, 3,650 executives and 3,350 employees: a sustained AI effect at 32% of companies.
  8. The GenAI Divide: State of AI in Business 2025 - the primary source for the figure of 95% of pilots with no measurable effect.
  9. The Simple Macroeconomics of AI - Daron Acemoglu, estimates of AI’s effect on productivity.
  10. Block lays off 4,000 of 10,000 staff, citing AI - the case engineers read.
  11. Scrum isn’t needed: how AI coding is reshaping the idea of an effective team - why traditional practices break and what replaces them.
  12. Konstantin Doronin: a framework for rolling out AI coding to a team - start with support tasks, cap pull-request size.
  13. The “LLM under the hood” channel chat: a staged rollout strategy - pilot, full coverage of one team, pitch the metrics, roll out.