Is Prompt Work Worth It If the Next Model Makes It Obsolete?
Judge prompt work by what it describes and whether it pays back before it expires. Preserve useful definitions of success and make model-specific instructions cheap to replace.
Is Prompt Work Worth It If the Next Model Makes It Obsolete?
When I used GPT-3.5, asking the model to work through a problem step by step and showing it one example of the response I wanted helped compensate for capabilities it lacked. These were chain-of-thought instructions and one-shot examples. As models improved, I began to question those habits. An instruction that once helped could become redundant, or constrain a model that had learned a better way to do the work.
That raises a practical question: if the next model may make your prompt work obsolete, how much effort should you put into it now?
Summary
Prompt work can be worth doing even if the next model makes it obsolete. The benefit it earns before expiry can repay the effort you invested. Two questions help you decide how much effort to spend:
- Does it describe a target or a path? A target states what counts as a good result; a path tells the agent how to get there. My hypothesis is that target information tends to survive model changes better than instructions adapted to one model.
- Will it pay back before it expires? Compare the effort of creating and maintaining it with the time, quality, or other benefit it returns while it remains useful. A durable artifact can be too costly for how little you use it, while a temporary one can repay its cost quickly.
Invest in reusable requirements where they repay their cost, and keep methods tailored to one model inexpensive to replace. These questions apply both when creating an artifact and when deciding whether to keep refining it.
This is a practical framework drawn from experience, with a durability hypothesis you can check against your own model changes. It is for individuals and small teams using AI coding agents, and concerns the artifacts you keep: project instruction files such as CLAUDE.md, reusable task instructions such as skills, automatic actions triggered during agent work such as hooks, and prompt templates. Take the task and your goals as given; the decision here is how much work to invest in helping the agent carry it out.
Yes, much of it expires
A compensation technique gets its value from a gap in the model's capabilities. If that gap closes, the technique loses its reason to exist. A detailed reasoning script may have helped a weaker model stay on track. A stronger model may already handle the problem, while the old script forces it through an unnecessary sequence.
There is a concrete, limited example of this change. OpenAI's guidance for its reasoning models says that explicit instructions to reason step by step may add no benefit and can sometimes hinder performance. The same guidance still allows examples when a task needs them. The lesson is to reassess an instruction when the model changes. OpenAI reasoning best practices.
Capabilities can also move into the model itself. On September 12, 2024, OpenAI released o1-preview, describing models trained to spend more time thinking before responding. That puts part of the reasoning work inside the model, reducing the need to prescribe it from outside. Introducing OpenAI o1-preview.
My experience has made me cautious about investing heavily in techniques for gaps many users share: a vendor may address the gap in a later model or product update. For an investment decision, the useful question is what the technique earns while you still need it.
Expiring is not the same as wasted
For an illustration, imagine that your coding agent sometimes reports a task finished without running the relevant tests. You spend an afternoon writing a hook that runs those tests before the task can be marked complete. For three months, avoiding the resulting debugging saves you half an hour each working day, after accounting for the extra test runs. Then the agent gains a native feature that reliably performs the same check, and you retire the hook. It has already earned back its cost.
Its expiry does not erase the time it already saved. The unrecovered part of the investment is what matters. A temporary artifact can be a good investment if its useful life is long enough to repay the work of creating and maintaining it.
So the original question opens into two others. How long is this artifact likely to remain useful? And how much benefit will it produce during that time? The first question helps estimate the window available for the second.
First question: does it describe a target or a path?
A target describes what counts as a good result. Acceptance criteria, tests, evaluation sets, business constraints, and examples of good output belong here. They carry information about what you want from the work.
A path describes how an executor should reach that result. Prescribed reasoning steps, role instructions, tool-call sequences, review gates, and divisions of work between agents belong here. Their usefulness depends more directly on the executor's current abilities.
Consider an illustrative import tool. Its target is that processing the same input twice must leave only one copy of each record. A path instruction for building it might require the coding agent to draft a plan, send that plan to a second agent for review, revise it, and only then implement the change. The requirement about duplicate records can stay valid when you change models. The mandatory planning and review sequence may need revision if the new model can reliably handle the work with less coordination.
The distinction also applies to examples. A sample showing one record after importing the same input twice demonstrates the behavior you want: a target. A worked example telling the agent exactly how to plan and delegate the implementation demonstrates a path. Both can help, but they carry different kinds of information. A file can contain both.
My hypothesis, based on observation, is that target content tends to survive model changes better than path instructions. The desired result can remain stable while the best way to produce it changes. You can check this against your last model upgrade: which requirements stayed intact, and which instructions did you have to rewrite or remove? If your target content also needed extensive revision each time, this distinction may predict little about durability in your setting.
An evaluation set is a collection of representative tasks with expected outputs or criteria for judging success. For the import tool, it could include a repeated input and a check that no duplicate records remain. You can keep judging that outcome while changing the model or its instructions. This is how evaluation sets make a target usable in practice, as discussed in Why Evals Are the Bottleneck for Useful Agent Systems.
DSPy, a framework for building and optimizing programs that use language models, provides a useful design example. Its prompt optimizers take the task program, examples, and a metric—a rule for scoring the output—and search for better instructions or demonstrations. The scoring rule tells you what result you value; the prompt is something you can adjust against it. This illustrates how to separate a definition of success from the wording used to obtain it. DSPy Optimizers.
Second question: does it pay for itself before it expires?
Durability supplies an estimate of useful life. It does not settle whether an artifact deserves your effort. A lasting specification that you use once may repay very little. A temporary workflow used every day may repay its cost quickly.
Compare the effort of establishing and maintaining the artifact with the benefit it produces before it stops working. Include the time spent debugging it, adapting it after updates, and carrying out any extra steps it imposes. On the benefit side, count what matters to your task: time saved, fewer mistakes, better output, or some other result you value.
You do not need to reduce every benefit to minutes. A quality improvement can justify work even when it saves no time. But the improvement needs to be something you care about enough to pay the cost. A longer prompt, an extra reviewer, or a more elaborate skill is worthwhile only through what it helps you achieve.
The comparison is especially useful when deciding whether to refine an existing artifact. The time already spent is gone. Estimate whether the next round of work will return enough benefit during the useful life that remains.
There is no general deadline that settles this for every model and project. Use your own experience of updates and usage. If a newer model can cheaply help rewrite the instructions, include that lower replacement cost too. Cheap rewriting makes temporary prompt work easier to justify.
Four kinds of prompt work
Putting the two questions together gives four ways to treat an artifact. The examples below are illustrations; their placement depends on the actual costs and benefits in your project.
| What it describes | Pays back before expiry? | How to treat it | Illustration |
|---|---|---|---|
| Target | Yes | Invest in an explicit, reusable asset | Acceptance criteria and tests used throughout a repository |
| Target | No | Reduce the work or skip the elaborate artifact | A full reusable evaluation suite for a migration script you will run once |
| Path | Yes | Build it so you can discard or replace it cheaply | A review sequence or agent division of work that helps on the current project |
| Path | No | Stop investing in or maintaining it | A large library of reasoning and role templates whose upkeep exceeds its benefit |
The second row is an easy one to miss. Even an artifact that describes the right target can be overengineered. For a one-time task, a few direct checks may be enough; a reusable evaluation system needs a reason to justify its additional cost.
Apply the framework to what an artifact contains, rather than to its filename. A CLAUDE.md file can mix durable requirements with temporary workarounds. A skill can carry both a definition of success and a method that needs replacing. Those parts can merit different treatment.
What about the workflow you built?
This is where expiry feels most expensive. You may have invested in skills, hooks, review gates, and several agents working together. Those arrangements can be useful, yet still become poorly matched to a new capability level.
Return to the agent that sometimes skips tests. You might also require it to write a checklist and hand its work to a separate reviewer. If a later model reliably carries out the checks itself, those reminders and handoffs may consume time without improving the result. The tests still define an acceptable outcome; the process for making the agent run them needs to fit its current capabilities.
A hypothetical Windows development team at Microsoft illustrates the same point. Imagine its unusually skilled early programmers know the code and share its design assumptions. They can work with terse explanations and considerable individual discretion. As the team expands to include programmers with less experience or less knowledge of that code, the same approach can become hard to maintain: new members need decisions explained, clearer handovers, and more explicit review.
In this thought experiment, the software still needs to be maintained, but the capabilities and knowledge of the people doing the work have changed. The process needs to change with them. For AI agents, the change can run in the direction of increasing capability: a checklist or review sequence designed for a weaker model may become unnecessarily restrictive for a stronger one. In both cases, a workflow can have been useful and still need redesign when the executor changes.
Use the payback question here too. A workflow that prevents costly mistakes or saves repeated effort can repay its cost while it fits. Build enough of it to get that benefit, with a cheap way to retire the parts that become unnecessary.
Some process serves coordination between people. For an illustration, suppose two maintainers must agree before changing a public interface used by their customers. A model that writes flawless code still cannot supply their agreement about the intended change. That approval step may remain useful as models improve.
Ask: if the team consisted only of you, with the model remaining exactly as it is now, would this process still be needed? In the two-maintainer example, the need to obtain the other maintainer's agreement disappears, pointing to a human coordination role. A review that catches mistakes the current model still makes has a different reason to remain. Reassess that part when the model's behavior changes. Keeping the model fixed in this question helps isolate what the other people contribute.
Tuning words for one model
Repeatedly changing wording to get one model version to behave is a particularly fragile form of prompt work. Its result can be tightly coupled to that model, while leaving you with little additional information about the project or the desired output.
If a small adjustment immediately saves substantial work, it can repay its cost. Beyond that, continued polishing needs a clear payoff. Before another round of phrasing experiments, ask what benefit the next adjustment is likely to return while that version still matters.
Make rewriting cheap
The two questions suggest four practical ways to write artifacts you can afford to update.
- Store the target separately from the path. Keep acceptance criteria, constraints, tests, and evaluation cases in a place that remains useful when you replace the operating instructions. You should be able to remove a workaround without losing the requirement it was meant to satisfy.
- Keep temporary instructions short. Write enough to make the current workflow effective. Each additional rule has to earn both its reading cost and its future maintenance. Compact instructions are easier to inspect, replace, and compare with a simpler version.
- Use examples of good output to express judgment. A good example can communicate preferences that would take many rules to spell out. Preserve it for the target information it carries; if its usefulness depends on prescribing an old method, treat that portion as a path instruction.
- Make optional process easy to switch off. A review sequence, hook, or agent arrangement should have a clear way to be disabled or replaced when it stops helping. Keep the acceptance checks available so you can judge the result of that change.
What can accumulate through these revisions is explicit knowledge of the project: what it should do, which outcomes are acceptable, and how to recognize them. Much of that knowledge already has familiar forms: specifications, tests, and evaluation sets. When those artifacts repay their cost, they give the next model a useful starting point.
Takeaway
The possibility of expiry is a reason to control how much you spend on a particular version. Invest when the benefits you expect during its useful life justify the cost. Put reusable definitions of success into specifications, tests, and evaluation sets, and keep the method for one model inexpensive to replace.
Before your next prompt improvement, ask whether you are describing a target or a path, and whether the additional work will pay back before it expires. Those two questions give you a reason to invest, a reason to stop, and a way to carry useful work forward.