// writing
not even wrong
An employee gets a message from her boss and can’t work out what it’s asking. She suspects it’s AI-written, so she pastes it into her own AI, which explains it and offers to draft the reply. Afterward she summed up the working relationship: “I can’t crack the code of working with [my boss], because it’s just his AI and my AI going back and forth.”1
That’s one desk. The measured version of the same story comes from a survey of knowledge workers: the more people trust the model, the less they think about what passes through it, and the thinking that remains shifts to checking outputs.2 That’s understanding, delegated at scale, from both ends at once.
The documents are multiplying, and everyone feels productive. What nobody can say anymore is who holds what the documents say.
the hive mind has a name
A company’s knowledge feels like a hive mind because it works like one, and the folk intuition has a 40-year-old scholarly name. Daniel Wegner called it transactive memory, and he found it first in couples: two people who each remember their own domain and reliably know what the other knows.3
transactive memory noun
The memory a group holds that no single member does: each head keeps its own specialty, plus a working map of who holds the rest. Knowing the answer, and knowing who knows the answer, spread across people and functioning as one memory.
It’s real and it’s trainable. Teams that learn a task together outperform teams of individually trained strangers, and the advantage comes from the shared encoding, not from familiarity or liking. But the system has joints, and the joints are exactly where documentation lives. Writing knowledge down only moves it if someone absorbs the reasoning on the other end: the people who helped produce a document read it the way it was meant, and the people who only received it mostly don’t.4 An artifact that never gets retrieved into a head is just storage.
One caveat, carried as a feature: nobody ever validated transactive memory at the scale of a whole company, even between humans.3 The loop is being installed into a system whose company-sized version was never shown to exist. And the loop routes around both joints at once: one model writes without a person internalizing, another model summarizes without a person decoding. The who-knows-what map keeps updating anyway. It increasingly points at nobody.
the producers and the bill
Who’s filling the pile? Two producers, and only one of them is the one you’re picturing. The first is the confident tourist. Dunning-Kruger survives its statistical critics only as a direction, and the direction is enough: the people with the least grip on a domain overestimate their standing in it the most.5 What tourists lacked before was throughput; producing a long, structured document used to require knowing something. Now the AI-era version of the effect has been measured directly: give people a chatbot on a reasoning task and the classic pattern vanishes, because everyone miscalibrates, gaining about three points of performance and about four points of self-assessment, with the most AI-fluent users the worst calibrated of all.6 The tool doesn’t democratize competence. It democratizes the feeling of it.
The second producer is underwater rather than overconfident. When the researchers who named workslop, AI-generated work content that masquerades as a contribution, went back months later to ask who makes it and why, the recipe they found was managerial: unclear AI mandates plus overwhelmed teams.7 The tourist and the professional ship the same fluent nothing for different reasons, and the receiver can’t tell them apart.
What receiving it costs has been measured properly only outside the office. The maintainer of curl watched confirmed vulnerability reports drop from better than 15% of submissions to under 5% when the LLM wave hit, and killed a bug bounty that had run six years; in all that time, the fire-and-forget AI channel produced zero real vulnerabilities, while the same tools in the hands of one accountable human produced roughly 150 fixed bugs. Clarkesworld closed submissions when machine-written stories hit 500 in a month against 700 human ones. And the Linux kernel’s arc is the one to memorize: this spring the AI-written reports turned genuinely good, and the maintainers’ review burden rose anyway, because the cost driver is volume through human judgment, not the quality of the output.8
Inside companies, the only number anyone has is a vendor survey in which workers self-report roughly two hours of cleanup per incident, published by a sponsor that sells the coaching to fix it.7 The extrapolation from bug trackers to meeting rooms is mine, and I’m making it.
Brandolini’s law, coined in a 2013 tweet and much older in spirit, prices refutation at an order of magnitude more energy than production. The half everyone forgets is social: refuting a colleague’s document means telling a colleague, sometimes a senior one, that their contribution is empty.9 Most people pay neither cost. In the workslop survey, recipients quietly rated senders as less capable and less trustworthy, and a third started avoiding them. Nobody flags the slop. It gets filed, and the who-knows-what map silently reroutes around one more person.
the best case
The case that the model belongs in this loop is real, and it deserves its full weight. Its best study put an assistant trained on top performers’ conversations onto a customer-support floor: productivity rose 15% on average, the newest agents gained most, and some of the gain persisted when the tool was off. That’s the optimistic version of everything here, the organization’s tacit knowledge extracted from its experts and dealt to its novices in real time.10 Consultants with GPT-4 did better work faster on tasks inside the model’s competence; professionals wrote 40% faster with output graded higher; and a meta-analysis across 106 experiments found real synergy on content-creation tasks whenever the human alone was better than the model alone.11 On the right task, inside the frontier, with a person still judging, the loop adds. Conceded in full.
Now read the same studies for what they measured. The quality gains are clarity gains: Microsoft’s own controlled test found assisted emails rated 18% clearer by a blind panel, with no statistically significant difference in accuracy. Clearer, faster, prettier, and not one point more correct, in the sponsor’s own data.12 The mechanism behind the leveling is compression: novices rise toward the house pattern, experts gain nothing or slip a little, and the collective diversity of ideas drops by about 11%. That’s redistribution of knowledge the organization already had, with nothing new entering the stock, and the learning that should mint the next experts doesn’t show up: model-assisted students scored best while the tool was on and learned no more than anyone else.13
And the boundary conditions from the meta-analysis read like a description of the documentation loop, written as a warning. The gains sit in content creation; judging what a colleague sent you is a decision task, the quadrant where human-plus-AI loses. The gains sit inside the model’s frontier; the consultants in the field experiment couldn’t tell where the frontier was, and outside it their accuracy fell off a cliff.11 The gains require a human still validating before anything ships; that’s the node the loop automates away, on both ends. The best case for AI at work is a description of the loop with a person still in it.
what the dashboard says
The feeling arrives first. Output is the thing a dashboard can see, and output explodes: assisted emails run 167% longer than human ones without carrying more, and in the study that bothered to check, most people who’d just written an essay with a model couldn’t quote a line from it minutes later.14 The cost lands elsewhere, on the receiving calendar, two unlogged hours at a time, visible on no chart. Volume up, comprehension unmeasured: the quarter reads as a good quarter.
Then the knowledge stock stalls while the document count climbs. Clarity without accuracy, leveling without learning, diversity down. And the alarm that should trip is wired to the wrong signal, because offloading inflates your sense of what you know more reliably than it changes what you know: searching makes people feel smarter about things they never looked up, and across the replication literature, that self-perception effect is the one that holds while the memory effects wobble.15 The heads empty quietly, feeling full the whole way.
Here’s the strange part. The instrument for checking exists. A corpus method published last year detects LLM-assisted writing at population scale with under 3.3% error, and it’s been pointed at public text for two years: by late 2024, machine-assisted shares ran from about 10% of job postings to 24% of corporate press releases, and had stopped climbing.16 It has never been pointed inside a company. There is no published measurement, none, of how much of any organization’s internal writing is machine-made, or of what the people who filed it still know. The loop closes exactly where nobody’s looking, and the feeling of productivity persists because nothing is positioned to contradict it. The web’s version of this loop at least has a market selling fresh human data back into the pipeline. A company’s only fresh supply is its own people’s attention, and that’s being spent on triage.
A lot of my job is reading what models wrote, mine and other people’s, so take this last part as testimony rather than measurement. The documents I’ve learned to dread aren’t the wrong ones. Wrong is workable: a wrong plan makes a claim, a claim can be tested, and the test teaches somebody something. The kind I dread is fluent, long, structured, confident, dense with decisions, and after an afternoon inside one I still can’t tell you what problem it’s trying to solve, which alternatives were considered, or which of its choices a person actually made. There’s nothing to refute. There’s fog to clear, and fog-clearing is invisible work that costs real hours and makes you the difficult one for asking what the pages are for.
Physics keeps an old insult for work like that, handed down through a colleague’s memoir of Pauli: not even wrong.17 Too unmoored to argue with, which is worse than mistaken, because argument is how a group of people comes to know anything. Everything is documented now. Ask the dashboard who understands it.
Footnotes
-
Jacqueline Munis, “‘It’s just his AI and my AI going back and forth’”, Fortune (July 2026; first published March 2026). The employee told the story to a Skillsoft executive who coins “social offloading” in the piece; Skillsoft sells communication coaching, so the framing has an interest. The quote is the employee’s. ↩
-
Lee et al., CHI 2025 (Microsoft Research): 319 knowledge workers, 936 first-hand examples. Higher confidence in generative AI correlates with less critical thinking; higher confidence in one’s own expertise with more; the remaining effort shifts to verification, integration, and stewardship. Self-report, correlational, and a finding that runs against the sponsor’s product. ↩
-
Wegner’s transactive memory (1985), validated at team scale: Hollingshead’s couples experiments; Moreland’s radio-assembly studies, where the gain came from joint training rather than familiarity; Lewis’s 2003 measurement scale. The team-to-organization extension is listed among the field’s unresolved problems in Ren & Argote’s integrative review of the literature’s first 25 years. ↩ ↩2
-
Internalization is the explicit-to-tacit step in Nonaka’s SECI model of knowledge creation. Decodification: Hall, “Knowledge management and the limits of knowledge codification”, Journal of Knowledge Management (2006): people present at the codification interpret the codes more similarly than people who only receive them. The storage distinction is Walsh & Ungson’s retrieval condition on organizational memory (1991). ↩
-
Kruger & Dunning (1999) for the pattern. The mechanism (“incompetence blinds you to itself”) is contested by at least four statistical-artifact accounts (Gignac & Zajenkowski 2020; Nuhfer et al. 2017; Krueger & Mueller 2002; Magnus & Peresetsky 2022) and by McIntosh et al.’s trial-level test (2022), which found metacognitive efficiency flat across ability. The direction replicates; the essay uses only the direction. ↩
-
Fernandes et al., Computers in Human Behavior 175:108779 (2026): N=246 and N=452, LSAT-style reasoning with ChatGPT. Performance rose about 3.5 points against the no-AI norm; self-assessment rose more; the paper’s own section title is “AI use cancels the Dunning-Kruger effect”; and self-rated AI literacy predicted worse metacognitive accuracy, not better. A separate 2025 result (Tully, Longoni & Appel, Journal of Marketing) finds lower AI literacy predicts higher receptivity to AI; that one is about adopting the tool, not judging your own output, and the two don’t combine. ↩
-
The coinage and the numbers: BetterUp Labs and Stanford Social Media Lab, “AI-Generated Workslop Is Destroying Productivity”, HBR (Sept 2025): 40% of 1,150 surveyed US desk workers received it in the prior month, 1 hour 56 minutes self-reported per incident, $186 per worker per month. Self-report on both time and salary, fielded Aug-Sep 2025, no pre-AI baseline, and BetterUp sells the coaching; their own pages restate the time figure as 1h51m and “2 hrs”. The producer side: the Jan 2026 follow-up, “the result of unclear AI mandates and overwhelmed teams”, prescribing review processes that reinforce rather than offload human judgment. ↩ ↩2
-
curl: Daniel Stenberg, “The end of the curl bug bounty” (Jan 2026); 87 confirmed vulnerabilities and $100,000+ paid across the program’s life. The accountable-human counter-case is Joshua Rogers’s ZeroPath work, AI scanners plus maintainer triage, roughly 20% false positives, ~150 bugs fixed. Clarkesworld: Neil Clarke’s own count (Feb 2023); his 2025 update says the second wave is harder to detect, not smaller. Kernel: Greg Kroah-Hartman via The Register (March 2026): “Now we have real reports… it’s more stuff we have to review.” ↩
-
Alberto Brandolini, January 2013, on Twitter: “The amount of energy needed to refute bullshit is an order of magnitude bigger than that needed to produce it.” A folk aphorism with precedents in Swift and Bastiat, not a measured law. The social-cost reading and the recipient statistics (42% saw senders as less trustworthy, roughly a third less willing to work with them again) are from the workslop survey, same grading as above. ↩
-
Brynjolfsson, Li & Raymond, “Generative AI at Work”, Quarterly Journal of Economics (2025): 15% average productivity gain, concentrated among less-experienced agents; the most experienced saw small speed gains and small quality declines; gains largest on rare problems; measurable residual learning. ↩
-
Consultants: Dell’Acqua et al., “Navigating the Jagged Technological Frontier” (BCG field experiment): ~40% higher rated quality inside the frontier, bottom-half performers +43% against +17% for the top half; outside the frontier, correctness fell from 84.5% (no AI) to 60-70% (with AI). Writing: Noy & Zhang, Science (2023): 40% faster, 18% higher graded quality, low performers gained most. Meta-analysis: Vaccaro, Almaatouq & Malone, Nature Human Behaviour (2024): 106 studies, 370 effect sizes, studies from 2020-2023; on average human-AI combinations underperformed the best of human or AI alone (g = -0.23); synergy appeared on content-creation tasks when the human alone beat the AI alone (g = 0.46), and losses on decision tasks and when the AI alone was better (g = -0.54). ↩ ↩2
-
Microsoft WorkLab’s Copilot study: a 62-person blind panel rated assisted emails 18% clearer and 19% more concise, and “across all three tasks, there was no statistically significant difference in accuracy.” A vendor studying its own product, which is exactly why the accuracy null is usable. ↩
-
The diversity figure: Dell’Acqua et al., a 10.7% reduction in the collective diversity of produced ideas. The learning null: Fan et al., British Journal of Educational Technology (2025): ChatGPT-assisted students produced the best essays and showed no greater knowledge gain or transfer, a pattern the authors call metacognitive laziness. ↩
-
Emails: Li et al., WebSci ‘25: AI-generated workplace emails averaged 166.8% more words than human ones (193.4 vs 72.5), more formal and structurally complex, with no matching gain in substance or personalization. A controlled comparison against synthetically generated replies, not observed office mail. Essay recall: Kosmyna et al., MIT Media Lab (preprint, n=54, about 18 per group): 15 of 18 LLM-assisted writers couldn’t quote from the essay they had just written. ↩
-
Searching: Fisher, Goddu & Keil, JEP: General (2015), nine experiments; the inflation held even for unrelated questions and failed searches. The replication picture: Gong & Yang’s meta-analysis (2024), 22 studies, ~31,000 participants: the cognitive self-esteem effect is significant at d = 0.91 while objective memory-performance effects are not, and the famous “Google effect” priming result failed direct replication twice. ↩
-
Liang et al., “The widespread adoption of large language model-assisted writing across society”, Patterns (2025): distributional quantification with population-level error under 3.3%; by late 2024, LLM-assisted shares were 17.7% of consumer complaints, up to 23.8% of corporate press releases, about 10% of job postings, 13.7% of UN press releases, with growth stabilized. The method’s scope is public text by construction. Checked July 2026: no published corpus-level measurement of internal-only workplace text exists; workplace-analytics telemetry measures tool usage, a different thing. ↩
-
“Das ist nicht nur nicht richtig; es ist nicht einmal falsch.” Reported by Rudolf Peierls in his 1960 Royal Society memoir of Pauli and a 1992 Physics Today letter; an anecdote with one named witness, not a dated quotation tied to a named paper, and this footnote is the sourcing. ↩