Recovering German text from an 85-year-old Enigma message sounds like a big breakthrough. In this case, it was only the beginning.
A September 14 to 15 investigation used GPT-6 Astra and several specialist AI agents to recover MVUEH, an 82-letter German Army message from July 10, 1941. The agents helped review old documents, build search programs, test possible answers, and challenge the results.
Getting a coherent German translation was not enough to call the message solved. The recovered Enigma key also had to decode all 82 reconstructed letters and match information recorded separately from the main message. Once those checks lined up, the researcher published the code and search materials so others could examine the work.
The original document was difficult to read
MVUEH was German Army message No. 172. Its recovered text roughly translates to: “Please specify the route of march. I am in Rosenow, Rosenow. Immediate reply by radio.” The apparent signature, “Waschbusch,” remains uncertain.
Enigma itself was broken decades ago by Polish and Allied cryptanalysts. This project focused on one message that had remained unresolved because researchers did not know the exact machine settings used to encrypt it.
The surviving copy also created problems. Several handwritten characters were hard to read even with a published transcription available, so the investigation identified 12 positions where more than one letter was possible. Combined, those uncertainties created 13,824 possible versions of the source text.
Researchers recorded those possibilities before finding the answer. Doing so reduced the chance that they would reinterpret an unclear letter simply because another choice produced a more convincing sentence. For AI-assisted research, this enables models to generate plausible answers even when the underlying evidence remains uncertain.
A 14-letter clue helped rule out bad settings
A previously solved message contained the repeated place name ROSENOWROSENOW. Researchers suspected the same 14 letters appeared in MVUEH, which gave them a known phrase to test against possible Enigma settings.
In cryptography, a suspected piece of the original message is called a crib. Here, the clue eliminated settings that could not produce ROSENOWROSENOW.
The remaining 68 letters still had to produce coherent text, so a setting could not be accepted just because it matched the guessed place name.
Researchers also had another piece of information recorded separately from the message body. The message header contained data operators used to set up the Enigma machine before encrypting the main text. Since the original search didn't use the header, researchers could use it afterward as an independent check on the recovered settings.
Across 14,829,646 possible Enigma setting combinations tested against the header, only 923 produced the expected result. The header did not identify the final key on its own, but it eliminated most alternatives.
Finding the same answer twice was not enough
The researchers ran several additional checks after recovering the key. Separate programs produced the same message and header result, while another verification system matched the machine behavior tested in 42 sample cases.
Using the published transcription, a second search also found the same key, although it reused a phrase discovered during the first recovery. The project treats the result as confirmation of the earlier finding rather than an independent second discovery.
Re-encryption provided another check by testing whether the recovered settings could turn the decrypted text back into the original encrypted text. Passing that test showed the calculation worked in both directions, but it did not prove the researchers had found the correct historical interpretation.
To make a stronger case, the recovered text had to remain coherent beyond the 14-letter clue, while the message header also had to match the recovered settings. Several independent pieces of evidence therefore pointed to the same answer.
External review has begun as well, with SWARM publishing its own review on September 17 using a separate simulator. It reproduced the expected message and header result, although its authors did not repeat the full search, so the review confirms only part of the investigation.
The AI agents were researchers, not referees
Several specialist agents worked on different parts of the project, including source review and search development. A coordinating agent compared their findings and sent disagreements back for further checking.
Although the agents handled different parts of the work, the researcher still decided which source readings could be used and which hypotheses were worth testing. Agreement among the agents did not count as proof, since several models can reach the same conclusion from the same mistaken assumption.
For anyone experimenting with AI agents, the MVUEH case shows why agreement between agents needs an external check. The project made its assumptions visible enough for others to compare the result against the ciphertext, the message header, the recorded letter options, and the published code.
By publishing its assumptions, code, and verification materials, the project helped researchers to test the result independently. AI helped one researcher attempt more of the investigation, while the final answer still depended on evidence that could be examined beyond the AI conversation.