In London, a small AI startup is trying to teach a machine something more valuable than how to write code.
It wants to teach the machine how an AI lab thinks.
Inherent, founded by former Google researchers Edward Hughes and Louis Kirsch, is building an AI system called Faraday. According to The New York Times’ reporting on the company, Faraday collects a remarkably intimate record of the lab’s work: emails, instant messages, meeting notes, and researchers’ conversations with Faraday itself. Inherent then uses that record as part of the process of making Faraday better.
That makes Faraday interesting for a reason that has little to do with the sci-fi version of an AI suddenly rewriting itself into a superintelligence.
The lab itself is becoming part of the dataset.
And that points to a much nearer version of “self-improving AI”: companies making the accumulated judgment of their researchers legible to machines, then using those machines to automate more of the research process that created them.
If that model works, one of the most valuable datasets in AI may eventually be something companies already own access to: the working memory of their own organizations.
The lab is becoming the training set
Most conversations about AI training data start with the internet: books, websites, code repositories, images, videos, and other material produced before a model was trained.
Inherent is chasing something different.
Its public manifesto describes the company itself as part of a “recursive collective self-improvement” system. Research discussions, resource decisions, hardware, training data, experiments, and AI systems are all pieces of the loop.
Inherent’s stated vision is deliberately human-centered. The company argues that AI’s role in science should involve collaboration rather than simply automating scientists away, and says it wants human-machine systems that keep people involved in discovery.
The New York Times reporting shows what that idea looks like in practice. Faraday does not only see the final code or the successful experiment. It can see traces of the process that produced them.
That distinction matters.
A rejected hypothesis contains information. So does a debugging trail. So does a researcher explaining why an experiment failed, a meeting where one project gets killed in favor of another, or a conversation where an experienced scientist spots something a junior researcher misses.
Traditional organizations lose plenty of that knowledge. It disappears into old Slack threads, undocumented decisions, forgotten experiments, and eventually the heads of people who leave.
An AI-native research lab can try to capture much more of it.
The result is a different kind of institutional memory: one designed to be searched, reused, evaluated, and potentially turned back into machine capability.
“Self-improving AI” covers a lot of ground
This is where the terminology gets messy.
Recursive self-improvement, or RSI, can describe several very different things.
At one end, an AI model writes code that humans use to build the next model. Move up another rung and agents can run experiments themselves. Go further and they begin choosing which experiments to run. Beyond that sits the much stronger scenario: an AI autonomously designs, trains, evaluates, and improves a successor system, which then becomes even better at repeating the process.
The final version is the “liftoff” scenario behind many of the industry’s most dramatic predictions.
Public evidence currently supports the earlier rungs much more strongly than the last one.
Anthropic, for example, says in its recent analysis of AI-assisted AI development that more than 80% of the code merged into its codebase was authored by Claude as of May 2026.
Its engineers were also merging roughly eight times as many lines of code per day in the second quarter of 2026 as they were in 2024.
But Anthropic itself attaches an important warning to that number: more code is not the same thing as eight times more productivity. Code volume measures quantity, not necessarily value.
The research evidence tells a similarly nuanced story.
In one Anthropic experiment, Claude-powered agents worked on an open-ended AI safety problem and recovered 97% of the gap between a weak and stronger model, compared with roughly 23% for two human researchers. But humans still chose the problem and created the scoring system, the agents consumed about 800 cumulative hours of work, and the result did not transfer cleanly to production-scale models.
The doing is getting dramatically easier to automate.
Choosing what is worth doing remains harder.
That distinction is one reason METR avoids using RSI as a precise technical label. Its researchers argue that the useful question is whether the feedback from better AI into better AI development becomes strong enough to sustain its own acceleration.
They do not think the evidence settles that question. Compute, data, inference capacity, experimentation, and research-specific capabilities can all become bottlenecks.
So the industry currently sits in an awkward middle ground.
AI is already helping to build AI.
AI building its own successors without meaningful human direction remains a different threshold entirely.
The workplace-data question arrives much earlier
Inherent makes another piece of this story harder to ignore.
If the quality of an AI research system depends partly on observing how expert humans work, then workplace activity becomes more than an operational record. It becomes an input into the system itself.
That raises questions that have almost nothing to do with whether superintelligence arrives in five years, 20 years, or ever.
What exactly does Faraday ingest?
Can researchers exclude particular conversations?
How long is the underlying material retained?
Is the system trained directly on communications, retrieving them when needed, learning distilled lessons from them, or using some combination of those approaches?
What happens when a researcher leaves?
And who controls the value created from years of accumulated decisions, experiments, failures, insights, and discussions?
The available public reporting does not answer those questions in detail.
That is important because emails and workplace messages are not sterile technical telemetry. They can contain personal context, disagreements, mistakes, confidential information, and details never written with the expectation that they would become inputs into an AI research system.
If Faraday’s use of workplace communications falls within UK rules governing worker monitoring or related employee-data processing, current Information Commissioner’s Office guidance raises questions about lawful basis, transparency, data minimization, security, retention, and data protection impact assessments.
That guidance does not establish that Inherent has violated any rule, nor does the public information available establish exactly how those rules apply to Faraday’s architecture.
It does show why the data question deserves to sit near the center of the self-improving-AI debate rather than in the footnotes.
The better these systems become at learning from organizations, the more important it becomes to know what the organization allowed them to learn from.
Human judgment is still the bottleneck
For now, frontier labs keep arriving at roughly the same remaining advantage for humans: judgment.
Anthropic’s own evidence says AI systems are increasingly good at implementing ideas, running experiments, debugging systems, and optimizing against defined goals.
Humans still hold more of the responsibility for deciding which goals matter.
That is sometimes described as “research taste”—the ability to decide which question is interesting, which result is trustworthy, which apparent breakthrough is a dead end, and when the scoring system itself is measuring the wrong thing.
Inherent is effectively betting that some of this judgment can also become learnable.
That helps explain why its collection of organizational data matters so much.
Successful papers capture the result, but much of what researchers call scientific taste lives in the path that never reaches publication: abandoned experiments, arguments, revisions, intuitions, near-misses, and the reasoning behind a decision.
Capturing that process could make research agents significantly more useful.
It could also move another piece of expert work from something a person knows into something the organization can preserve in software.
That shift has economic consequences long before anyone builds superintelligence.
If a firm can convert more of its experts’ accumulated working knowledge into reusable machine capability, documentation, traces, experiments, and decisions can persist as machine-accessible institutional memory even after the employee who generated them is gone.
That changes what it means for a company to “know” something.
Oversight has to follow the feedback loop
The AI industry is already beginning to grapple with one part of this problem.
Anthropic CEO Dario Amodei recently proposed giving independent evaluators ongoing, employee-like access to frontier labs so they can inspect models, training processes, safety practices, and incidents rather than relying entirely on what companies choose to disclose publicly.
As The Neuron recently examined in the debate over slowing frontier AI, the proposal is partly an attempt to make safety commitments verifiable.
Systems like Faraday suggest that such oversight may eventually need to examine more than model behavior.
Evaluators may need to understand the feedback loop itself.
What data feeds the research agent? Who has access to it? What counts as meaningful human oversight? What happens when an agent produces more work than people can realistically inspect? Which internal metrics actually demonstrate research progress rather than more code, more experiments, or faster iteration?
And when a company says its AI is becoming better at improving AI, who gets to verify what “better” actually means?
Those questions become especially important when the evidence comes from the same organizations racing to develop the technology.
Independent evaluation is only as useful as the independence, access, and publication rights behind it.
The first feedback loop is already here
The public evidence stops well short of a system autonomously designing, training, evaluating, and deploying a superior successor.
The nearer transformation is easier to see.
AI labs are redesigning themselves so that more of the research process can be observed by machines, performed by machines, and fed back into the next round of research.
That loop does not need to become an intelligence explosion to matter.
It only needs to make the organization better at turning human judgment into reusable machine capability.
Inherent’s experiment makes the tradeoff unusually visible. The same emails, discussions, failed experiments, and small acts of expertise that make a research organization valuable can also become raw material for AI systems capable of performing more of the research process themselves.
So one of the first governance questions around self-improving AI may arrive well before anyone needs an emergency button for superintelligence.
It starts with something much more ordinary:
When a machine learns how an organization thinks by watching its people work, what exactly have those people agreed to help build?