AI Can Do It for You—But Should You Let It?
Practical Wisdom in the Age of Savant Machines
ChatGPT can now do most of your digital work. Should you let it? The answer is complicated because modern LLM-powered tools are like savant machines—systems with extraordinary abilities but limited wisdom. On the one hand, they can speed you up tremendously; sometimes, they seem to give you wings. On the other hand, they often produce polished answers that can conceal absurd failures. Every unchecked answer remains a gamble. Even if they were perfectly reliable, delegating to them can mean surrendering some of the practice that built your competence and judgment in the first place. Savant machines are too useful to ignore—but too uneven to trust with everything. So we need a way to decide when to use them. My current rule of thumb: delegate when the work is mechanical, good enough is enough, or the terrain is too broad to inspect alone. Retain judgment and do the work yourself when the goal is unclear, the stakes are high, or the struggle is part of the point. Use AI to extend your thinking, not to replace it. Savant machines may transform the world, but they do not understand what it means to live in it.
Over the last few months—and especially the last few weeks—the economics of my work have fundamentally changed. My days as a software engineer and writer look nothing like they did before. Early AI chatbots often felt like unreliable autocomplete and were easy to dismiss. For my work, however, the tools of 2026 have crossed an important threshold.1 They now take a credible swing at most digital tasks I throw at them. On an ordinary Thursday afternoon, I can summon a small herd of agentic minions, set them loose on the mechanical work in front of me, and watch a full day’s work tumble into an hour or less. In those moments, I feel far more capable than I have any right to be.
Speed is only the most obvious way these tools extend my capabilities. What matters more to me is that they collapse the cost of curiosity.2 Until recently, every one of my side quests competed with my existing workload. Often, an idea could be genuinely captivating and still too expensive to pursue. Now, instead of settling for the first reasonable option, I can explore seven. I can follow digressions, test half-formed ideas, and run experiments I might never have considered at all. Over the last two weeks alone, I’ve developed four ideas into full-fledged apps simply because it has become cheap enough to find out whether they are useful. I also wrote at least four different versions of this essay before publishing this one.
But it also feels a bit like black magic. The same agent that compresses a day’s work into an hour can just as easily consume the next three. Yes, I built four apps. Yes, I wrote several versions of this essay. But not all that exploration was useful. I spent more than three times as long on this essay as I usually do. Some of those hours went into reading more deeply and pursuing ideas I would never have explored without AI. Others disappeared into refinement loops that left me unsure whether the problem was the model, the argument, or my own judgment. I often could not tell where curiosity ended and pointless drift began until much later. No other technology I have used has exerted such a strong pull on my attention.
The confidence these tools give me is real. With their help, almost any problem feels at least approachable. While writing this essay, they helped me find missing ideas, sharpen awkward phrases, and expose weak arguments. At other times, however, I asked for help before I knew what I wanted to say. The machine returned smooth, plausible sentences, and I mistook forward motion for progress. Several paragraphs later, I realized I had followed an argument I did not actually believe. That is the darker side of the bargain. Every answer remains a gamble. LLMs still invent facts, hide weak assumptions inside polished prose, and carry an argument forward long after it has stopped making sense. As the tools improve—and as I become better at using them—the odds shift further in my favor. But their failures also become easier to miss. Like inverse slot machines that almost always pay out, they train me through repeated success to lower my guard. The failure that matters then arrives precisely when I have stopped expecting one.
But my greatest fear is no longer that these tools might do my thinking badly without me realizing it. The longer I use them, the more I worry about something deeper. Using AI too much—or in the wrong ways—may not merely give me bad answers. It may erode the practice on which my own thinking and judgment depend. A bad answer can often be corrected. Becoming less able to recognize one is much harder to notice. For all their magic, these tools increasingly reach beyond knowledge work and into my mindwork itself. They tempt me to hand over precisely the difficult thinking through which I learned to judge their output in the first place. Over time, that may weaken the expertise on which both my livelihood and my professional identity depend.
Over the last few weeks, I therefore kept returning to two questions: What should I delegate to LLMs? And what should I insist on doing myself—perhaps precisely because a machine could do it for me? This essay is my attempt at an answer. I won’t pretend there is a clean rule. After all, the same work may be safe to outsource when the stakes are low and the result is easy to verify, yet worth doing myself when the struggle is part of how I learn. But I think we can find heuristics that get us most of the way. To find them, we first need to look inside these strange machines: how they work, why their impressive abilities break down in such peculiar ways, and where their limits lie. Then we will turn the lens back on ourselves and ask the more unsettling question: What happens to us when we let savant machines think on our behalf?
I. Savant Machines
Large language models are hard to understand and even harder to explain. Ideally, anyone who relies on them would learn the basics of how they work.3 But even among those who make the effort, it is almost impossible to talk about LLMs without metaphor. In the introduction alone, I have already compared them to minions, inverse slot machines, and black magic. Others call them copilots to emphasize assistance, stochastic parrots to warn that fluency can emerge from probabilistic pattern-matching without understanding, or blurry JPEGs of the web to describe their lossy compression of what humans have written online. Each metaphor isolates something real. But none quite captures the contradiction that matters most to me: radical unevenness.
My preferred metaphor for this comes from an old medical label. In 1887, the British physician J. Langdon Down used the phrase idiot savant to describe people who displayed extraordinary ability in a narrow domain alongside significant intellectual limitations. The phrase is now obsolete and offensive; the current term is savant syndrome.4 What remains fascinating is the radical unevenness: extraordinary ability in one domain can coexist with substantial limitations elsewhere. That same unevenness makes me think of large language models as savant machines—artificial systems in which a “savant” and an “idiot” share the same architecture. Such a machine can perform astonishingly well in one moment and stumble over something elementary in the next.
Their unevenness can be almost comically easy to see. Models have famously miscounted the r’s in strawberry, claimed that 9.11 was larger than 9.9, and insisted that it’s better to walk to a nearby car wash. However, the latest LLM-powered tools hide this contradiction much better. ChatGPT, Claude, and their peers can now count, write, code, summarize, search, and carry out tasks across many fields, often with astonishing skill—or so it seems. From the outside, then, they appear broadly capable—true generalists—which raises the obvious objection: in what sense are they still savant machines?
The answer is that breadth is not the same as evenness. A language model can operate across many fields while remaining radically uneven within them. When it summarizes a book, it can cover every chapter while missing the book’s central argument or fall prey to the well-documented lost-in-the-middle effect, overlooking relevant material simply because it appears far from the beginning or end. It can also find and cite a real legal case but attribute to it a convincing quotation that never appeared in the opinion. And, of course, it can write faulty computer code that passes all the tests it wrote, because it didn’t write that one test case that actually mattered. In each case, the model gets enough right to conceal what it got wrong. The pieces fit together, but they form the wrong picture.
Knowing whether the picture is right requires more than seeing that the pieces fit. It requires judgment and discernment. The problem is that language models do not reliably signal when they are wrong or their confidence is misplaced. A weak answer can sound just as convincing as a brilliant one. The tools of 2026 no longer tell me to walk to a nearby car wash. But while editing this essay, they vehemently insisted that I add the explanatory clause “forgetting that the car itself needs to come along” to the example—as if a human reader would otherwise miss the joke.
So that is what I mean by a savant machine: a large language model with extraordinary abilities but limited wisdom, capable of producing both brilliant and absurd outputs, yet lacking the self-reflection or reliable signals that would tell you which is which.
II. Borrowed Judgment
If a savant machine cannot reliably distinguish its brilliant outputs from its absurd ones, how do systems built around it become more reliable over time? After all, the progress itself is clearly measurable.5 The question, then, is not whether the system has improved. It is where this newfound judgment and growing reliability come from.
Part of the answer is scale. During training, a base language model acquires much of its competence indirectly, through artifacts created by humans. A research paper may condense years of work into a few pages. A mature codebase may embody thousands of failures found and fixed. These artifacts also preserve human error, fashion, and shallow imitation. But at their best, they carry the residue of experiments the model never conducted and failures it never lived through. A model can reproduce the shape of expert work without retracing the process that made the work trustworthy. In this limited sense, LLMs borrow condensed human knowledge—and some of the judgment embedded within it. As models grow and are trained with more data and compute, their language predictions improve in a remarkably regular way. Better data and training methods can push their reliability further.
But scale can take a model only so far. Its built-in knowledge stops at its training cutoff, and some tasks—such as precise calculation—remain unreliable when left to the model alone. That is why many modern AI systems add a layer of reliability outside the model itself. They combine the model with retrieval, external tools, memory, and execution loops. The resulting system can call a calculator, search for current information, retrieve documents, test results against clear criteria, and revise failed attempts. The savant machine has not overcome every weakness. The scaffolding around it catches or works around many of them and is often called an agent harness. So when you use one of these modern tools, you are not dealing with a savant machine alone. You are dealing with a scaffolded savant machine (SSM): a model embedded in a larger computer program with explicit instructions, access to tools, and checks on its work.
Both the savant machine and its scaffolding are shaped by human knowledge and choices. People selected, filtered, and assembled training data. They wrote demonstrations, ranked model outputs, and used those rankings to train reward models. They specified principles intended to shape model behavior, designed evaluations, chose tools, wrote tests, and decided what counted as success. Humans observed the world, conducted experiments, and lived through the consequences. The machines borrow much of that judgment. So what looks like the LLM itself becoming more reliable is often its scaffolding at work. Each layer expands the range of tasks the system can perform dependably by supplying information, tools, or checks the model itself lacks. Retrieval supplies information, calculators supply computation, memory supplies context, and execution loops supply further attempts. The scaffolding makes the savant machine far more reliable. This added reliability, however, depends on someone having defined success. Scaffolding works best against failures people anticipated and knew how to catch. When that borrowed judgment runs thin, the model may keep producing beautiful next steps toward nowhere in particular.
III. Where Borrowed Judgment Ends
Even the latest SSMs remain unreliable in important ways. Only a few weeks ago, I gave Claude Code a medium-sized code file and asked it to find problems with it. It confidently returned six. I replied to each finding individually with the same sentence: “This is wrong.” Six times, Claude explained why its original finding had been mistaken. This experiment proved neither the findings nor the retractions; it showed that I could steer the model’s evaluation with a single confident sentence, and it could make either position sound convincing. Researchers call it sycophancy when a model sides with the user instead of giving the answer more likely to be true. That differs from a hallucination, where a model makes up something plausible and states it as fact. In one case, the model bends toward me; in the other, it drifts away from the evidence. But both leave me with the same problem: neither the fluency nor the confidence of an SSM’s output tells me whether it deserves my trust.
Both failures can be reduced—or at least caught more often—through defenses at both layers: human judgment frozen into the model through post-training on human demonstrations, labels, and preferences; and evidence, tools, tests, and review loops supplied at runtime through the harness. But these defenses reach their limit when deciding what counts as success is itself part of the problem. An SSM can test whether a feature works, but it cannot decide for us whether the feature should exist in the first place. It can check whether evidence supports a claim, but it cannot decide for us whether publishing that claim would be fair to the people involved. It can compare two career options against your stated goals, but it cannot decide how much time with your family you should sacrifice for ambition. The first question in each pair comes with a standard; the second asks us to choose one. A supported claim may still cause harm. Ambition may pull against family, or one person’s interests against another’s. Before a benchmark can score the answer, someone must decide what success means, which costs are acceptable, and who will have to live with them.
The kind of judgment needed here has an old name. Aristotle called it phronesis, or practical wisdom: the ability to judge well about what to do in a particular situation. Ethical virtue points us toward worthwhile ends; practical wisdom helps us decide what they require here and now. Aristotle contrasted it with cleverness, the ability to find effective ways of reaching a chosen end, whether that end is good or bad. An SSM resembles cleverness in this limited sense. Give it a target, and it can help you find and follow a route toward it. It can even help you examine the target. But competence alone cannot make the target worth pursuing or relieve us of responsibility for choosing it.
Seen this way, it helps to distinguish at least three capacities when we talk about AI: Competence is the ability to carry out a given task. Discernment is the ability to judge whether the work is sound. Practical wisdom is the ability to decide what is worth doing in a particular situation—and to take responsibility for what follows. SSMs of 2026 are already remarkably competent across many tasks, which is why I sometimes have those lucky days when a full day’s work tumbles into an hour. With increasingly powerful scaffolding, they are also getting better at low-level discernment wherever humans have supplied a clear standard or encoded their judgment. But practical wisdom is a different game. Here, an SSM can still support us by revealing options, challenging assumptions, and surfacing relevant precedents. But it cannot exercise practical wisdom for you. After all, practical wisdom is more than producing good advice. It is exercised by someone who must choose, act, see what follows, and live with the result. An SSM has no family to neglect, no career to risk, and no life that can go better or worse because of the choice. No amount of training or scaffolding changes that. It can improve the counsel, but it cannot assume the stakes of the choice. Ultimately, the call—and responsibility for it—remains yours.
IV. Scaffolding Amnesia
So we have built this scaffolding around a savant machine. When we use the resulting system, we draw on knowledge acquired and judgment exercised elsewhere. We no longer have to acquire all that knowledge—or exercise all that judgment—ourselves. But what happens when borrowing these capacities means using our own less often? What does that do to our competence, discernment, and especially our practical wisdom?
Software has long altered our habits and patterns of thought. But most conventional tools enter after an intention has begun to form. A chat interface or semi-autonomous SSM can enter much earlier, while the problem is still unclear and our own position has not yet taken shape. It can give us an explanation before we understand the subject, an argument before we know what we think, or working code before we understand how or why it works. For a while now, I have watched one reflex grow stronger in me: when difficult mental work arises—an unfamiliar idea to grasp, an argument to shape, or a messy thought to untangle—my first instinct is no longer to sit with it. It is to open the chat.
Opening the chat also changes the shape of my attention. Once several agents are running, sustained concentration begins to feel inefficient: every minute I spend thinking deeply about one problem is a minute I am not using to steer the others. I am increasingly inclined to move from agent to agent—prompting, reviewing, correcting, and starting new tasks—rather than staying with one difficult question. More work moves forward, but I do less of the sustained thinking through which difficult problems become clear.
None of this felt like dependence at first. It felt like efficiency. Why struggle for an hour when the machine can produce an answer in seconds? Why move one problem forward when I could move five? But both questions hide the same assumption: that only the visible result matters. In mindwork, that assumption is often false. Mental effort does more than produce the answer in front of us. It also builds the vocabulary, background knowledge, experience, taste, and mental models through which we understand a subject and judge future answers. The result may still be perfectly useful. But when learning, growth, and skill-building are part of the goal, some of the work that produces them must remain ours.
A scaffolded savant machine can extend—and sometimes replace—my thinking in the moment while weakening the habits that sustain it over time. Difficulty, feedback, mistakes, and repetition are part of how we develop our own judgment. Handing over the work can therefore mean giving up one of the reps through which that judgment develops. Yet none of this is obvious in the moment. We see the completed work, not the practice we skipped. That makes us far too willing to outsource our thinking without considering what repeated delegation may do to us over time. Much as social media reshaped our attention through repeated use, AI may gradually reshape our capacity to think for ourselves.
I call this scaffolding amnesia: forgetting that yesterday’s struggles are a big part of the reason we are where we are today. It applies to both the machine’s scaffolding and our own. It makes us overestimate the machine’s judgment while underestimating how much practice our own judgment still requires. The output may remain strong even as the person behind it changes. But experts may gradually lose abilities they already possess, and beginners may never develop them at all.6
V. When to Reach for the Machine
I know of no good way to draw a precise boundary in advance between what to delegate and what to keep. Everyone probably has to cross it a few times before they learn where it lies. Still, some heuristics can help us decide when to reach for the machine—and when to keep the work for ourselves.
When You Feel Like a Monkey
Reach for the machine when you sense there must be a better way—one with less typing, copying, pasting, and reformatting. Reach for it when you catch yourself repeating the same small task for the tenth time, or whenever your hands are busy but your mind is barely involved. When you feel less like a thinking person and more like a trained monkey pulling levers, you have probably found work worth delegating.
This is the natural home of automation: repetition, grunt work, and the operational glue between systems. You already know what the result should look like, which rules it must follow, and what counts as good enough. You have properly engaged with the work, and the path is clear. Only the execution remains. At that point, your fingers are the bottleneck.
I saw this recently when I decided to improve and streamline my weekly reviews. Reviewing the past week and planning the next is invaluable; clicking through my Obsidian vault is not. So I automated many of the manual steps, cutting the time they take in half. I now spend the remaining 30 minutes almost entirely on judgment calls. Scripts handle the rest.
💡 Maintainable Artifact over Instruction
Once a workflow becomes stable and repeatable, turn it into a durable artifact rather than asking an LLM to perform the same steps again and again. My weekly-review improvements live in scripts. My productivity middleware lives in an app. These artifacts run faster and behave more predictably. One caveat: be honest and intentional about the dependency you add to your life. At least for now, a script you cannot fix without an LLM subscription can quickly turn from an asset into a recurring liability. So, whenever possible, prefer artifacts you can maintain by hand.
When Good Enough Is Enough
Some work is low-stakes, one-off, and not worth perfecting or automating. It still needs to be done, but it does not deserve craftsmanship. This is Pareto work: an 80-percent solution is enough because the final 20 percent would cost more than it is worth.
LLMs are useful here. They can help you build a rough prototype, fix a low-risk bug, draft a thoughtful email that does not need to be literature, or create a tiny app that only needs to solve today’s problem. In many of these cases, the alternative is not a better result but no result at all. The idea stays untested, the annoyance unfixed, or the temporary tool unbuilt because none of them is worth several hours of careful work. The SSM lowers the cost of trying.
When I recently used Claude Code to rebuild my Squarespace website from scratch, not every detail needed to be perfect. I mainly wanted to preserve the content and keep roughly the same look and feel. Without AI, the project might have taken me ages because I have little experience with web design—and I probably would have skipped it entirely. With AI, I could reach a working version, put it online, and move on. As web design is not part of my craft, I didn't need to understand every line, and the code didn't need to be beautiful. It only needed to work well enough for as long as it was meant to exist — Pareto work for me.
💡 What counts as good enough?
Perfectionism predates AI, and AI can make it even worse. When another iteration costs only another prompt, the final 20 percent always looks affordable. Cheap iteration widens the search when you are exploring, but pulls you deeper into the same groove when you are polishing. Satisficing is part of practical wisdom. It means choosing the standard the situation actually requires. A disposable script may be ugly. A prototype may have rough edges. But legal documents, medical decisions, financial calculations, and public claims leave far less room for error. Pareto work is only Pareto work when the remaining flaws are cheap, visible, and reversible.
When You Need to Survey the Terrain
LLMs are especially useful when the evidence is scattered across more material than you can comfortably hold in your head. An SSM can search notes, transcripts, books, health logs, or research papers and surface recurring themes, related passages, and possible connections. Its advantage here is breadth, not depth. You define the question and the boundaries; the machine performs the wide, tedious first pass. The key here is that you treat the result as a map of where to look next, not a verdict.7
I used AI this way during a recent health investigation. After years of gaining weight, standard tests had explained little, and the usual advice rarely went beyond sleeping better, eating less, and exercising more. None of that was necessarily wrong, but it did not tell me whether something more specific was worth investigating. So I asked ChatGPT to help me survey the terrain—not to diagnose me, but to map potentially relevant blood markers, what they might indicate, which tests overlapped, and how I could assemble a focused panel within my budget. The broader panel then indeed surfaced findings that earlier testing and conversations had not brought together for me. ChatGPT did not diagnose me, and the results still required professional interpretation. Its contribution was narrower but real: it helped me compare more information than I could comfortably review alone and turn a vague concern into concrete evidence and better questions.8
💡 Make the search checkable
Used well, the SSM gathers and arranges the evidence. But you remain responsible for deciding what it means—often with qualified help. Give the machine a narrow task and clear criteria. Ask it to quote the passages behind each claim, link to the original sources, distinguish supported findings from plausible interpretations, identify exceptions, and mark open questions. Then verify anything consequential yourself or with someone qualified to do so. Let it make your thinking wider and more demanding—not merely faster and easier.
VI. In Defense of Formative Difficulty
The three cases above share an important condition: the standard remains yours—or belongs to whoever is accountable for the outcome. With mechanical work, you know the desired result; with Pareto work, you decide what counts as good enough; with breadth, you supply the question and interpret the evidence. In each case, the machine extends your capabilities without taking the standard away from you.
Sometimes, however, forming that judgment is the work. I call this formative difficulty: difficulty that matters partly because working through it develops knowledge, skill, taste, or discernment that outlasts the immediate task. Not every difficulty qualifies. Much of it is accidental: clicking, copying, reformatting, and other monkey work. Remove that without regret. But when I am learning a field, forming a position, or developing taste, the process is part of the result. The work is changing me as well as producing an answer.
This is where the ease an SSM provides becomes risky. We may mistake the clarity of its answer for progress in our own understanding. Confusion and mental discomfort can signal that our first frame does not fit, an assumption is weak, or the real question lies elsewhere. Staying with them gives us a chance to discover what is wrong. Ian Leslie describes a related local-minimum problem: an LLM can optimize impressively within the frame supplied by a prompt while leaving the frame itself untouched. It can suggest alternative framings when asked, but it does not reliably recognize when the original question should be abandoned or reframed. If we reach for the chat before examining that tension, we may get an answer before we understand the problem: borrowed clarity without the judgment needed to know whether the answer deserves our trust.
Learning science gives me reason to take this concern seriously. Psychologist Robert Bjork calls certain forms of productive struggle desirable difficulties: conditions that impair performance in the moment but improve long-term learning. A related finding, the generation effect, suggests that we often remember material better when we generate it ourselves than when we simply receive it. That fits my experience: I remember the essays I wrote before ChatGPT better than those I wrote with AI assistance. Neither line of research directly tests AI. But together they suggest that some struggle does more than delay an answer: it helps us become able to produce—and judge—the answer ourselves.
Early AI-specific evidence points in the same direction. A 2026 preprint reporting a randomized experiment found weaker conceptual understanding, code reading, and debugging among developers who used AI while learning a new programming library. Those who fully delegated the coding gained some productivity but learned less. Yet learning was preserved among participants who asked conceptual questions, requested explanations alongside generated code, or checked their understanding afterward. This is one study in one domain, not evidence of general cognitive decline. But it sharpens the trade-off: the important line may run not between using AI and avoiding it, but between replacing engagement and sustaining it.
Tim Ferriss once posed one of modern productivity’s most useful questions: What would this look like if it were easy? It remains a useful question, but an SSM can remove both accidental and formative difficulty without telling us which one disappeared. The question is not wrong now, only incomplete: So I pair it with another: What makes this hard—and which part of that difficulty is still mine to carry?
A Practical Test. Over the next few days, whenever you use a scaffolded savant machine, ask yourself:9
Am I delegating execution—or the decision about what matters?
Am I removing drudgery—or avoiding a formative struggle that could teach me something?
Do I understand the problem well enough to judge the result rather than merely accept it?
If the machine is wrong, will the failure be cheap, visible, and reversible?
Did this session sharpen my thinking—or merely leave me relieved that I no longer had to think?
VII. Scale the Monkey Work. Own What Matters.
Searching, summarizing, coding, writing, calculating, sorting, and any other kind of digital work—AI can now probably do it for you. The question is, should you let it? Turn the question around: What should remain ours? Not every difficult task. Not every act of thinking. But certainly enough of the difficult practice through which our competence and judgment are formed. Whatever we delegate, the final say over what matters—and responsibility for what follows—remains ours.
I do not yet know exactly where that boundary lies. By the time a loss of competence becomes obvious, the habits that caused it may already be deeply established. I do not want to wait for proof before becoming more deliberate about what I hand over.
Use the scaffolded savant machine to make execution easier, search more widely, and experiment more cheaply. But keep deciding what matters, judging what is true, and taking responsibility for what happens next. That is practical wisdom in the age of savant machines: accepting their help without surrendering the final call.
Scale the monkey work. Own what matters.
More specifically, I currently use GPT-5 models through ChatGPT Plus for writing. For coding and most other work, I mainly use Claude Code and Cowork with Opus and Fable on the Claude Max 20× plan. For smaller recurring tasks, I run qwen3.6:35b locally through Ollama within PiHarness.
The distance between wondering and trying has almost disappeared. Yet my to-do list has grown—a version of the Jevons paradox, in which greater efficiency increases rather than reduces total consumption. AI cheapens exploration and lets you clear work faster than ever, but it also multiplies what feels worth exploring. The only thing faster than AI, it seems, is curiosity.
At a minimum, I recommend watching Andrej Karpathy’s Deep Dive into LLMs like ChatGPT, a three-and-a-half-hour overview aimed at a general audience.
People with savant syndrome are often simply called savants. If you have never heard of this condition, here are two well-known examples to make the contrast vivid: Leslie Lemke reportedly played Tchaikovsky’s Piano Concerto No. 1 after hearing it once and yet struggled even to hold eating utensils. Kim Peek, who died in 2009, could memorize thousands of books yet could not button his own clothes. In both, extraordinary ability in one domain coexisted with severe limitations elsewhere.
METR’s software-heavy task-completion time-horizon benchmark shows that, compared with the systems METR measured in 2024, recent frontier agents have substantially longer time horizons. At the same success probability, they are predicted to complete well-specified tasks estimated to require substantially more time from a human expert.
Scale that dynamic from individuals to a culture, and scaffolding amnesia becomes more unsettling. Pixar’s 2008 film WALL·E imagines its physical logic: when machines remove the need to move, bodies lose the ability. Applied to the mind, that same logic points toward Mike Judge’s 2006 satire Idiocracy: a society in which convenience, confidence, and anti-intellectualism outlive the knowledge that once sustained it—a society that has forgotten how to think for itself.
But pay close attention to the filter it uses. Whatever the machine leaves out may never reach your attention. What you see can quickly become all that seems to exist—the AI version of WYSIATI: what you see is all there is. The aim is to outsource the volume of attention required—not the discernment that tells us what matters.
My bloodwork revealed very low vitamin D, severely elevated LDL cholesterol that was later traced to a genetic cause, and markedly elevated fasting insulin despite normal glucose and HbA1c. I’m now taking meds and supplements and, at least subjectively, feel a lot better.
Note that while these questions can sharpen your awareness, they do have a blind spot: they ask you to use your judgment to determine whether that judgment is weakening. If AI compensates for abilities you exercise less often, the combined system may keep producing excellent work while your unaided ability slips. The output may remain strong precisely because the machine is masking the change you are trying to detect.
So complement self-assessment with two tests:
Test the output. Require a relevant test to pass, inspect the underlying evidence, or ask another person to challenge your reasoning. These checks tell you whether the work currently holds up.
Test yourself. Occasionally exercise the ability you want to preserve without the machine: form the argument, solve the problem, or find the path on your own. This is not an exercise in virtue or technological purity. It is a diagnostic—a way to distinguish what you and the machine can accomplish together from what you can still do yourself.






