← Back to Notes

AI can rewrite your legacy system. Can you prove it's the same system?

Mistral's Fortran migration is being read as proof that agents made legacy modernisation cheap. What their report actually shows is that the translation was never the expensive part.

Mistral published an account recently of a legacy modernisation project for a European energy operator, and the version doing the rounds is that AI agents migrated Fortran to C++. They did, but that’s not what I found interesting.

The report is narrower and, to their credit, more candid than the summaries of it. The codebase was a physics-heavy reservoir simulator, around 300,000 lines of Fortran 77 with no test suite and no centralised documentation. 40,000 of those lines were migrated, which is impressive and also about 13% of the codebase. Neither the duration nor the cost of the migration was disclosed.

The first attempt turned the agents loose autonomously. What came back was Fortran retyped in C++ syntax with the old architecture intact - COMMON blocks became global structs and GOTO-driven control flow stayed as it was. Mistral’s own verdict is that the result was functional, “but it couldn’t be called code modernization”. The second attempt used a multi-agent pipeline, which improved matters and then hit a bug, tried a few fixes, and stalled with nobody there to unstick it.

What actually worked looks a lot like how a careful team would have done before “the AI age”.

Using the correct tooling

A custom parser modeled the procedural code as a caller-callee tree. Through this process they could pinpoint the modules that operated independently. They subsequently wrote subroutines that export the state of the running Fortran program at designated checkpoints. They also built a test framework that loads those checkpoints into the C++ code for comparison. The teams defined agreement as equality between the final results and the intermediate points that the client’s reservoir engineers had marked as critical.

It was at this stage that the agents proved valuable. Over one hundred agents, using Mistral’s Vibe CLI, documented the codebase beginning with the highest‐level components and proceeding downward. They processed paper documentation via optical character recognition (OCR). A human then coordinated a cycle of coder, tester, and reviewer agents, resolving issues and integrating the suggested changes. The reservoir engineers had already examined the planned architecture before any code was written.

Note the specific strengths of the agents. They were effective at producing documentation, mapping dependencies, and carrying out repetitive translation tasks within an established framework. Note the order of events.

You can’t test against something you can’t run

Each test needs an oracle to determine the correct result. When developing new software, the oracle is usually a documented specification, a product owner, or a decision made in a meeting. When porting existing software, the oracle can differ.

Twenty years of updates have replaced the original specification, and the original authors have all retired. Even if they had not the “parity harness” is going to be the most critical part of the process. This method requires executing the Fortran code in isolation with controlled inputs. Mistral describes the self‐contained and runnable nature of this codebase as a “favorable starting condition”. They also note that migrations face more challenges when they depend on external systems, lack a runnable baseline, or contain undocumented physics.

I think that is the most significant sentence in the article, yet nobody has quoted it. An agent can build a comparison harness for you in a single afternoon - however it cannot generate the required state to populate that harness. If the legacy system only runs within a mainframe partition shared with other processes, or if its behavior depends on a nightly data feed that cannot be replayed, or if the logic exists solely in the mind of one engineer, the agents will have no real baseline for comparison. They will still generate C++ code, and that code will compile. You will simply have no way to determine whether it really functions the same as the original system.

This kind of test isn’t a new idea

Mistral has built a particular type of test. In the 2004 book “Working Effectively with Legacy Code”, Michael Feathers defined them as characterisation tests. These tests record how the code currently functions, including any existing bugs, to identify when a change alters that behavior. The goal is not to verify that the code is correct. Instead, these tests provide a baseline for comparison with the new version.

A parity harness is a characterisation suite employed for simulators. It operates by exporting state, recording outputs, and comparing the results. Most legacy systems lack these suites because they were tedious to write and did not include new features. AI agents don’t care if the work is boring or not - so they have removed these obstacles.

A standard quote allocates most of the funds to the rewrite and a smaller portion to testing. If translation costs are minimal and the testing harness determines the project’s success, the budget should focus on archaeology and so that we can properly define an oracle. A bid that puts most of the spending in the translation phase indicates that the bidder has not yet examined the legacy system.

Feathers identified a specific problem with this method. A characterization suite records both the intended features and the existing bugs. If the system contains a rounding error from 2003 that downstream reports already account for, the test harness will treat that error as the correct behavior. A person with domain expertise must decide which behaviors are part of the system design and which are mistakes. AI agents can spot that two numbers differ, but they cannot decide which number is correct for that project.

What this means for large code estates

The National Audit Office reported in January 2025 that government departments were using at least 228 legacy IT systems as of March 2024. There were no fully funded plans to modernise 120 of these systems. These are the specific systems that are currently being considered for migration to AI.

Consider the definition of a runnable baseline in this context. The business rules consist of decades of legislation, ministerial amendments, and operational workarounds. These rules are not documented anywhere except within the code and the memories of the staff who have maintained it for decades. The systems exchange data batches with other legacy systems. Existing test environments are only partial copies of the production environment. Mistral identifies the problem of complex knowledge existing only in one engineer’s head as an additional challenge, but in a benefits system this is the standard condition. Core banking systems follow the same pattern. In those cases, the overnight settlement run serves as the actual specification. I doubt anybody has read the codebase in its entirety for years.

Agents are still useful in this context. They are most helpful for documentation and dependency mapping. On a project of this scale, identifying the existing technical structures is a task that can take several years. The most expensive aspects of the work remain the same. I previously stated that public sector IT is difficult because it must account for complex real‐world edge cases, and that the primary challenge is not the technology. Modernizing legacy systems involves the same principle. It requires determining how the current system affects the end user and ensuring the new system meets those same needs.

How I’d scope one of these projects

Begin by asking a specific question that people rarely raise during an initial meeting. Determine if you can run the existing system in isolation, provide controlled inputs, and record the resulting outputs. On most systems I have encountered, the honest answer is “not without a lot of work”, and that work constitutes the first phase. This stage contains all the uncertainty of the task and requires appropriate pricing and staffing. If the answer is truly yes, the conditions are favorable and the remaining work is standard engineering.

Build the harness before you translate anything. Use the agents for what they are demonstrably good at: reading the documentation that does exist, mapping the call graph, finding modules that stand alone, drafting the characterisation tests. Run many of them, because that work is cheap now and it is what always got cut in the past.

Decompose and migrate the system one module at a time where possible. A person must determine the meaning of “the same” whenever the harness identifies a discrepancy. Every difference falls into one of three categories: a bug in the new code, an existing bug in the old code that the business has accepted, or an undocumented rule that was previously unknown. This third category represents the primary value of the process. I believe the discrepancy log is a more useful deliverable than the code itself.

AI agents have reduced the cost of rewriting legacy systems. However, they have not reduced the cost of determining how those systems currently function. A rewrite remains a significant risk based on incomplete information until you understand the original logic.

Got something you'd like to talk through?

A short email is plenty. Tell me where you are and what you're wrestling with.