Store Production Prompts in Your Application Code
OpenAI's deprecation of its Prompts API—with fourteen words of migration guidance—forces a question most production teams are poorly equipped to answer: where do prompts live, and who owns their history? The directive to version them in application code is correct; the problem is that most organisations lack the evaluation discipline and audit infrastructure to comply without embedding new technical debt in the same location as the old.
OpenAI announced on 3 June 2026 that it will shut down v1/prompts on 30 November 2026. The migration guidance is fourteen words long. Keep each production prompt in a code-managed, versioned helper such as prompts/supportReply.ts. No replacement service, no hosted alternative, no tooling to smooth the path. The engineer reading that notice knows exactly what was cut to ship Frontier. The company that acquired Promptfoo in March does not have a first-party answer to the question it just created for everyone who built on top of the API it is now removing.
OpenAI forced API users to transition from Assistants to Prompts, and now they are taking away the Prompts. Developers have lived through the migrations from Assistants to Prompts to now. What the developer forum thread makes plain is that unless the instructions are placed within a database, any minor change to the prompt will require a rebuild and redeploy. The managed object gave product teams a way to iterate prompts without touching code. That workflow is gone. OpenAI is telling its API users to treat prompts as typed functions, version them in git, cover them with tests, and gate changes behind pull requests. It is prescribing exactly the discipline that most enterprises lack, because there is no equivalent yet of continuous integration or continuous delivery for prompts.
In practice, many temporary prompts become long-lived, mission-critical assets. Many teams still manage prompts as hardcoded strings buried inside application code, with no version history, no audit trail, and no way for non-engineers to iterate without triggering a deployment. The debt compounds silently. Different teams often implement similar AI use cases with slightly different prompts, and over time these prompts drift, leading to inconsistent tone, guidance, or recommendations across the organization. Nobody can see the divergence until a user complains that the AI behaves differently depending on where it is accessed, and by then the fix requires an archaeology project to reconstruct which version of which prompt landed in which service.
The companies that sell prompt versioning have known this was coming. In 2026, prompt versioning has evolved from basic version tracking to complete development infrastructure, and the best tools connect versioning to evaluation, enable staged deployment through environments, and provide collaborative workspaces where product managers and engineers iterate together. LangSmith launched in February 2024 with a per-seat model at $39/month and usage metering on top. By October 2025 LangChain closed a $125 million Series B led by IVP at a $1.25 billion valuation. The entire category went from niche to funded in under two years because the problem OpenAI just handed back to its users is the one every production LLM deployment hits. Actual LangSmith costs are 10.7x the base subscription when accounting for real-world implementation, archival, and operational overhead. The $39 seat price is only the beginning. Base traces run $2.50 per 1,000 with 14-day retention, and extended traces cost $5.00 per 1,000 with 400-day retention once you exceed your included quota.
The adjacent market is moving faster than the tooling can keep up. 98% of FinOps teams now manage AI spend, up from 31% in 2024, but 73% of AI projects still blow budget. Average monthly AI spend hit $62,964 in 2024, with projections showing it will rise to $85,521 in 2025, a 36% year-over-year jump. The number one most-requested capability across the entire survey is granular monitoring of AI spend, covering tokens, LLM requests, and GPU utilization. Commercial tooling has not delivered this at scale. The FinOps Foundation reclassified AI cost management as core scope in the span of 24 months, but the infrastructure to track what a given prompt costs per inference, per user, per session has not caught up. Prompt versioning sits upstream of cost allocation. If you cannot tie a trace back to a specific prompt version running in a specific environment under a specific feature flag, you cannot answer the question a CFO asks when the bill doubles.
The governance layer lags even further behind. General-purpose AI model obligations under the EU AI Act kicked in August 2025, and the penalties are real, up to €35 million or 7% of global turnover, whichever is higher. The companies treating governance as an afterthought are building technical debt that will cost them exponentially more to fix later. A regulator asking for an audit trail of how a decision was made does not care whether the prompt was managed in OpenAI's deprecated API or in a git repository. What matters is whether you can produce the exact version of the prompt that was live when the contested inference ran, plus the test coverage that gated its deployment, plus the evaluation scores that justified promoting it to production. A 2025 MIT study found that 95% of AI projects fail to reach production or deliver value, and a similar study by S&P Global Market Intelligence found that 42% of businesses scrapped multiple AI initiatives in 2025, a sharp increase from 17% the previous year. The denominator in both numbers is partially explained by teams that shipped a proof of concept, could not operationalize it, and abandoned the work when the cost of maintaining unversioned prompts across a changing model surface exceeded the value of the use case.
According to LangChain's State of AI Agents Report 2025, 57% of organizations have agents in production, 72% of enterprise AI projects involve multi-agent architectures, and 49% of enterprises have 10+ agents running in production. An agent system that chains three prompts and calls two tools generates traces with nested spans. Agents route between sub-agents, call tools, retrieve context, and chain reasoning steps, and when one of those decisions goes wrong the failure cascades through every step that follows. Prompt versioning for a single-turn assistant is a documentation problem. Prompt versioning for an agent that composes prompts dynamically across a state machine is an architectural dependency, and the fact that OpenAI removed the one hosted mechanism it offered tells you something about the maturity of the tooling required to do it well. Promptfoo, which OpenAI acquired in March, has attracted more than 350,000 developers and 130,000 active monthly users across multiple AI providers and models. Promptfoo's capabilities address the specific blockers CIOs are naming, including red-teaming, compliance monitoring, audit trails, and behavioral testing that convert blocked deployments into production workloads. The acquisition removed a barrier, but it did not add a feature to Frontier that replaces the API OpenAI is now deprecating.
The engineer reading the migration guide on 3 June knew three things immediately. First, that prompt management was never core to OpenAI's platform strategy, or it would not be the second feature to get deprecated in a migration cycle that already burned Assistants users. Second, that the directive to version prompts in code is correct advice and also advice that most product teams cannot follow without writing their own abstraction layer or buying a seat in a vendor's closed platform. Third, that the technical debt clock started the moment the first team hardcoded a prompt string into a Lambda function in 2023, and the compound interest is now coming due in the form of unevaluated agent sprawl, FinOps overruns, and compliance risk that nobody can quantify because the audit trail does not exist.
The deprecation notice is 72 days old. The shutdown date is 72 days away. What changes between now and then is not whether teams will move their prompts into code. They will, because the alternative is that the application stops working. What changes is whether they treat the forced migration as an opportunity to build the versioning, testing, and evaluation discipline that makes prompts legible to the rest of the organization, or whether they lift the strings into a constants file, commit it without review, and preserve the technical debt in a new location. The companies that finish the migration in the first way will have an artifact they can tie to a cost line, an evaluation score, and a compliance record. The ones that finish it in the second way will discover in Q2 2027 that they cannot explain to an auditor which version of which prompt produced the output now under review, and the bill for reconstructing that answer will be a multiple of the original deployment cost.
Tarry Singh is the founder and CEO of Real AI (realai.eu), an enterprise AI advisory and deployment firm working with global enterprises on production agent systems, model risk, and AI sovereignty strategy. He also leads Earthscan (earthscan.io) for Energy AI, and is a founding contributor to the EU-funded HCAIM and PANORAIMA programmes for responsible AI education across European universities. He writes at tarrysingh.com.