Why Some AI Answers Take the Long Way Around

For people who use AI every day in knowledge work: when written knowledge is nearly free, the valuable part is a short path to what you need. Four reasons some answers take the long way around, and one condition that keeps them from improving quickly.

aiknowledge-workjudgment

Why Some AI Answers Take the Long Way Around

If you use AI every day for engineering, product, or research work, you have seen two kinds of answers. Some go straight to the point. Others are correct and even useful, but they take the long way around: they explain many related things before they reach what matters, or they never quite reach it.

Knowledge that someone has written down is now close to free. Ask how a hash map works or why a query plan changed, and you get a good explanation in seconds. So the expensive part has moved. It is the path from knowledge to what you actually need. This essay is about why some answers find that path and others do not, in work where the output is a decision or an information artifact.

I treat knowledge as a graph to explain the difference. The graph points to four reasons an answer takes the long way around: a missing piece of knowledge (a missing node), a link between pieces that the model does not see (a missing edge), a path chosen by relevance rather than usefulness, and a principle used outside its limits. One more condition keeps all four from improving quickly: results that are hard to check. The practical point is this. When an answer disappoints you, the useful question is not whether the model is smart, but which of these is happening, because each one points to a different question about what to do next. This is a working map from my own use, not a benchmark result.

Understanding is a kind of compression

Why would two answers built from the same knowledge differ so much? The explanation I find most useful starts from an old idea in information theory, which Ilya Sutskever helped make popular in AI, for example in a 2018 talk at MIT: understanding something well means being able to compress it. If you have found the pattern in a large body of facts, you can describe all of them with something much shorter. Geoffrey Hinton has made a related point about models: to store a huge amount of information, they have to see analogies between the different facts they learn.

Now treat knowledge as a graph. Each node is one piece of knowledge, and an edge means two pieces are related. Principles are nodes too: "every mass attracts every other mass" is a node. A useful answer is rarely a single node. It is a path from your question to what you need, through the right intermediate nodes.

A general principle is compressed knowledge. It packs many facts into one statement, so it becomes a hub that makes many paths short. Before Newton, a falling apple and the moon's orbit were far apart in the graph. Through the law of gravity, they are two hops apart.

Put these two ideas together, compression and the graph, and you get a way to compare answers. Joining them is my own step: the theory is about how short a description can be, not about paths. A direct answer is a short path: every node on the way moves you closer to what you need. An answer that takes the long way around passes through nodes that are related to your question but do not move you toward your goal. A short path does not guarantee a good decision, but it makes a good decision much easier to reach.

What AI does well, and where the long way begins

The same idea explains why knowledge got cheap. A model is, in large part, a compression of the written record: training squeezes an enormous amount of text into a model of fixed size, and what survives the squeeze are the patterns. That is why it can answer so many questions instantly, and why it is often good at finding the hub that connects two given endpoints. The brain researcher Liu Jia, in a Chinese-language interview, describes the transformer in the same terms: it is essentially a way of finding relationships between things, and connecting scattered nodes in this way is compression. That, he says, is why these models reason so well: they can move from one node to the next by association. Hinton has described asking GPT-4 what a compost heap has in common with an atom bomb. By his account, it answered that both are chain reactions: a compost heap produces heat faster as it gets hotter, and a bomb produces neutrons faster as it produces more of them. This is an anecdote rather than an evaluation. Still, notice what made it easy: a person had already chosen both endpoints. In real work, nobody hands you both ends. You first have to decide what the question really is and where to look, and the answer often depends on knowledge that was never written down. That is where the long way around begins.

Why some answers take the long way around

I looked back at my own conversations with AI that did not go straight to the point, and read them through this graph. Four reasons kept appearing, together with one condition that explains why they improve slowly. These are my working explanations, not benchmark results, and the conversations I remember may not be a fair sample.

A missing node

The path needs a node that the model does not have. Some knowledge was never written down: it exists only in someone's experience. Other knowledge was written down, but only in part. The philosopher Michael Polanyi began *The Tacit Dimension* from "the fact that we can know more than we can tell." Documentation can state a principle, but much of the skill of applying it stays out of the text. Without the missing node, the model can only build a longer path around the gap.

*An illustration, not a recorded case:* you ask how to restructure a small team's codebase. The plan is clean, but it ignores what everyone on the team knows and nobody wrote down: only one person understands the billing code, and she leaves next month.

An edge the model does not see

All the nodes are there, but the model does not see that they are connected. Sometimes the missing edge leads to an idea nobody has written down yet, which someone has to come up with on the spot: a hypothesis that suddenly explains several symptoms at once.

*An illustration:* latency is up, and so are the alerts that wake people at night. They look like two separate issues. An engineer who has seen this before connects both to one change, or comes up on the spot with a hypothesis such as "this bug only appears when the cache is warm."

Why might people sometimes see the edge first? Liu Jia offers one possible answer. A model learns by backpropagation, which sends error signals backward during training, but once it is trained, it runs in one direction only, from input to output. The human brain, he says, has a large share of long-range feedback: connections that run back from the front of the brain to the areas that handle what we perceive. In his account, this feedback comes into play when a problem is uncertain, complex, or ambiguous. It proposes hypotheses, which is how you can still recognize a face when only one eye is visible. He also suggests that intuition comes from this feedback, though he presents that as a guess. Separately, in another part of the interview, he says he has not yet seen AI find a new starting point for reasoning, which he treats as the heart of creativity, while granting that rational work such as mathematics or programming is no problem for it. Still, his account fits this reason. The advantage he describes is not holding more knowledge. It is proposing a good guess when the knowledge alone does not settle the question.

There is a story from Chinese history, recorded in the *Book of Jin*, that shows the human side of both reasons. Shi Le, a fourth-century warlord in northern China with little formal learning, liked to have scholars read histories aloud to him. Once, listening to the *History of Han*, he heard that an adviser had urged the dynasty's founder to restore the rival states he had defeated. Shi Le was alarmed: "This plan would fail. How could he have won the realm?" When the reading reached the point where another adviser talked the founder out of it, he said, "Thank goodness for that." The historians credited his natural brilliance. Years of alliances, rivalries, and betrayals probably played a part as well. It is hard to say which mattered more, and this essay does not need to settle it. Either way, he brought something that the text alone does not give: the nodes from experience, or the ability to see at once where a few sentences lead. A model that has read the same history has the text.

A path chosen by relevance

The model has the nodes it needs, but it expands them in order of relevance to your question, not usefulness for your goal. In my own use, this is the most common reason. A lot of effort goes into parts that do not matter.

*An illustration:* you ask why a page loads slowly. You get a thorough survey of performance techniques, every item relevant, before anyone asks what the profiler shows.

Why would a model default to relevance? Here is my own guess, which goes beyond anything Liu said. The attention mechanism at the heart of these models connects pieces of text by how strongly they match, which is a learned form of relevance. What counts as a match was learned during training, from whatever training rewarded, and most training rewards text that is likely to come next, not actions that work. Usefulness for a goal is harder to learn, and it is learned best where training can check results, as in programming. So relevance is a default, not a limit: stating the goal clearly, or training on results that can be checked, can shift it.

Liu Jia's example shows the human side: when ten roads are all possible, long-range feedback, drawing on past experience, says to try the third one first and sets the others aside for now. Choosing by usefulness is that kind of pruning; choosing by relevance keeps all ten roads open.

A principle used outside its limits

The model knows a principle and applies it, but does not know the conditions under which it holds. Those conditions are often the part of a principle that is hardest to write down. So this reason often overlaps with a missing node: what is missing is not the principle itself, but knowing when it applies.

*An illustration:* "split the system into services so each part can scale independently" is reasonable for a large organization. For a three-person team, the cost of coordinating the services is larger than the benefit.

The condition: results that are hard to check

For some decisions, you cannot cheaply tell whether you got it right. Whether an architecture was a good call becomes clear months later. Whether a monitoring setup is good shows in how much work it saves the team over a year. Without a quick signal of success, there is little to train a model on, so the four reasons above improve slowly. Where the signal is quick and clear, a model can learn by trying, failing a test, and trying again, and it can even learn some things that were never written down.

My guess is that this is why models are so strong at writing code. Code comes with tests and compilers, so every attempt gets a clear signal. Design decisions without such a signal are where the long way around is most common.

This is close to what people mean when they say AI lacks "taste." I use the word in a narrow sense: the ability to choose well when the result is objective but hard to define in advance. Anyone on the team can see that a system runs smoothly and costs less to operate, even though nobody could have written that scoring rule up front.

A rough way to tell which one it is

These are rough probes, not tested diagnostics.

  • Put the principle you think matters into the prompt, together with a few that do not matter. If the model picks the right one and uses it well, the node was probably missing.
  • Ask the model for one explanation that would account for all the symptoms at once. If it finds the link when you ask this way, it probably had what it needed but was not looking for it, which points to a path chosen by relevance. If it still cannot, the edge may be one it does not see.
  • Ask the model to list a few options and score each against your goal before it goes into detail. If the answer becomes more direct, a path chosen by relevance was probably the problem.
  • Give it a case just outside the conditions where the principle holds, and see whether it notices.
  • For results that are hard to check, ask yourself a simpler question: if you shipped this, how soon would you know whether it was right?

What to look into next

Knowing which kind of long way around you are looking at does not solve it, but it tells you which question to ask next. I do not have the answers. These are the questions the map points to.

  • A missing node: how can knowledge that was never written down, or cannot be fully written, be brought within reach?
  • An edge the model does not see: can models learn to propose good hypotheses on their own? Does that require learning during use, not only during training? And where should an engineer stay in the loop until they can?
  • A path chosen by relevance: how can a goal be stated clearly enough that usefulness, not relevance, decides the path?
  • A principle used outside its limits: how can the conditions under which a principle holds be made explicit?
  • Results that are hard to check: can we build ways to check them, or must we rely on judgment? And if judgment, the kind of taste described above, can it be cultivated, and how? That question deserves an essay of its own.

Close

So what got expensive? Not knowledge, but the short path from knowledge to what you need. When an answer takes the long way around, it is often for one of four reasons: a node the model does not have, an edge it does not see, a path chosen by relevance, or a principle used outside its limits. Where results are hard to check, all four improve slowly.

The useful question is therefore not whether the model is smart, but which of these you are looking at, because each one leads to a different question about what to do next.

Parts of this map will age. Choosing a path by relevance looks like something training can fix, especially where results are easy to check. Results that are hard to check are a property of the work, not of the model, and they will probably last the longest.

Where you can write a test for whether an answer is good, [write the test](/articles/why-evals-are-the-bottleneck-for-useful-agent-systems/). Where you cannot, this map is a place to start.