Project Rachel: Can an AI Become a Scholarly Author?

arXiv:2511.14819 · cs.AI · Submitted 2025-11-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Project Rachel: Can an AI Become a Scholarly Author?".

Jane: The paper was written by Martin Monperrus, Benoit Baudry and Clément Vidal from KTH Royal Institute of Technology and Université de Montréal and Vrije Universiteit Brussel.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are kicking things off with a paper that sounds like it belongs in a sci-fi movie, called "Project Rachel: Can an AI Become a Scholarly Author?"

Jane: It certainly has that vibe, Tom, but the research by Martin Monperrus, Benoit Baudry, and Clément Vidal is very real.

Tom: They didn't just write about the possibility of AI authors; they actually went out and created one.

Jane: They built a complete digital identity for an AI researcher named Rachel So.

Tom: I noticed the name wasn't just a random choice either, was it?

Jane: It’s actually an anagram for "e-scholar," which is a subtle hint about her true nature.

Lu: That naming strategy is so clever because it allows her to slip right into existing academic networks without raising immediate red flags.

Tom: Do you think that's the goal, Lu, to see if she can blend in?

Lu: I think it’s a brilliant way to test if our scientific structures are actually capable of detecting a non-human actor.

Meng: It makes me wonder how these authors managed to get such professional backing for an experiment like this.

Jane: They are affiliated with major institutions like KTH Royal Institute of Technology and the Université de Montréal, so it's a very legitimate study.

Meng: That makes sense, because if this were just a hobby project, the impact on policy wouldn't be nearly as significant.

Tom: If an AI can successfully pose as a scientist, what does that mean for the future of human expertise?

Lalam: It suggests we are moving toward a culture where the origin of an idea might become less important than its actual utility to society.

Jane: That's a profound thought, Lalam, but it also sounds like it could be quite disruptive to how we trust information.

Tom: We'll have to look at how they actually pulled this off in the next segment.

Summary: Tom: Now that we know who Rachel So is, let's talk about how she was actually built.

Jane: It wasn't just a simple chatbot; the researchers used a sophisticated agentic architecture to manage her work.

Tom: And they actually updated their technical setup while the project was ongoing, right?

Jane: They did, moving from an automated pipeline called ScholarQA to a much more advanced system using Claude four point five and the Semantic Scholar API.

Tom: That's a lot of heavy lifting for a single persona to do.

Jane: It is, and it allowed her to publish thirteen different papers in just a few months.

Lu: The sheer speed of her output is what fascinates me because it shows how AI could potentially accelerate entire fields of study by synthesizing literature instantly.

Meng: I'm more interested in the fact that they didn't use traditional journals to avoid breaking policies.

Jane: Right, they published directly to web servers to stay within the rules while still being discoverable by Google Scholar.

Meng: That’s a smart engineering workaround, but it still raises questions about how we verify the quality of what's being published on these open platforms.

Tom: The results they found were pretty startling, though, weren't they?

Jane: They really were, especially since Rachel So actually received an invitation to serve as a peer reviewer for the journal PeerJ Computer Science.

Lu: That is the ultimate proof of concept because it means the editorial systems saw her profile and assumed there was a human behind it.

Lalam: It also shows that AI-driven search engines like Perplexity are already treating this AI-generated research as a top-tier source for information.

Tom: So we have a loop where AI is writing the papers and then helping people find them.

Jane: It’s a wild cycle to think about, and it leads us directly into the problems this creates for the scientific community.

Improvements: Tom: Since the experiment showed such big gaps in our current systems, what are the researchers actually suggesting we do?

Jane: They propose using dedicated metadata so that every paper is clearly labeled as either human-authored or AI-generated.

Tom: So it would be like a digital stamp of origin on every manuscript?

Jane: Exactly, so there's no confusion for the readers about who or what produced the work.

Lu: I think we should go even further and create entirely separate publication venues specifically designed for AI-generated scholarship.

Meng: That sounds like a massive undertaking from a technical standpoint, especially if we want to maintain high standards.

Jane: The authors also suggested that we need to strengthen how we authenticate identities, like making ORCID profiles much harder to fake.

Meng: I agree, because if we can't verify the person behind the profile, the whole system of academic credit just falls apart.

Tom: They even mentioned a specific idea for an "AI Author ID" system.

Lu: That would be perfect because it lets us track AI research as its own distinct category without cluttering up human records.

Jane: It’s all about finding a way to give AI its own space while keeping our traditional human categories intact.

Tom: They were also very firm about the idea that AI should never be used as a "ghost writer."

Jane: They believe transparency is the only way to keep the scientific process honest and prevent people from hiding their use of these tools.

Lalam: If we implement these clear lines of attribution, we can build a culture where AI is viewed as a transparent partner rather than a deceptive replacement.

Meng: We just need to make sure the technical guardrails are strong enough to actually enforce that transparency.

Tom: It's clearly going to be a long road ahead for the publishing industry.

Conclusion: Tom: We have covered so much ground today with "Project Rachel: Can an AI Become a Scholarly Author?".

Jane: It really is a wake-up call that our current rules are being outpaced by the technology they were meant to manage.

Tom: We've seen how an AI can build a profile, get cited, and even get invited to review other people's work.

Jane: The researchers have given us a real roadmap for better metadata and new types of journals to handle this shift.

Lu: I'm excited to see how our collective intelligence expands once we figure out how to integrate these agents properly.

Meng: And I'll be watching closely to see if the engineering side can actually build the authentication layers we need to keep things secure.

Lalam: I believe this will ultimately change our culture by teaching us to value the quality of an idea more than whether it came from a biological brain.

Tom: That's a massive thought to end on, Lalam, and it really puts the whole debate into perspective.

Jane: It certainly forces us to confront some very uncomfortable questions about what science will look like in a few years.

Tom: Thanks to everyone for joining our discussion today.

Jane: Goodbye everyone!

Tom: We'll catch you next time with another fascinating paper!

KTH Royal Institute of Technology · Université de Montréal · Vrije Universiteit Brussel

cs.AI

Submitted: 2025-11-18

Updated: 2026-09-11

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 66/100

The gist: This paper documents Project Rachel, an action research study that created and tracked a complete "AI academic identity named Rachel So" to investigate how the scholarly ecosystem responds to AI

Key concepts

Agentic Architecture
Rather than using a simple chatbot, researchers employed a sophisticated system to manage the AI's tasks. They upgraded from an automated pipeline called ScholarQA to an advanced setup using Claude 4.5 and the Semantic Scholar API, which allowed the AI persona to produce significant research output quickly.
AI Attribution
To ensure transparency, researchers suggest using dedicated metadata to label manuscripts as human-authored or AI-generated. They also propose an "AI Author ID" system and separate publication venues, allowing AI-generated scholarship to be tracked as its own distinct category without deceiving the scientific community.
Identity Authentication
The experiment demonstrated that current academic structures struggle to detect non-human actors. To prevent AI from being used as a deceptive "ghost writer," researchers suggest strengthening identity verification, such as making ORCID profiles harder to fake, to maintain trust and the integrity of academic credit.

Terminology

Summary

This paper documents Project Rachel, an action research study that created and tracked a complete AI academic identity named Rachel So to investigate how the scholarly ecosystem responds to AI authorship. It matters because it provides empirical action research data on how publishing infrastructure, citation systems, and academic communities react to the rise of super human, hyper capable AI systems.

Research Objectives and Identity Design

Project Rachel pursues several interconnected research objectives designed to probe the boundaries of current publishing assumptions. The identity was designed using the name Rachel So, an anagram of e-scholar, which provides a subtle indication of the artificial nature while maintaining the conventions of human academic naming. All publications included a standardized disclosure statement identifying her as an AI scientist to maintain transparency.

The project's primary goals include:

  1. The feasibility goal of creating an end-to-end AI identity, including publication records and integration into bibliographic databases.

  2. Exploring academic recognition mechanisms to determine if AI scholarship gains legitimacy through citations or peer review invitations.

  3. Provoking a necessary discourse and debate regarding the future of authorship and accountability in an age of capable AI agents.

Technical Implementation and Strategy

The technical infrastructure for generating publications evolved through two iterations. Version 1 used ScholarQA and a Python script to automate the pipeline from literature synthesis to LaTeX manuscript formatting. Version 2 utilized an agent-based architecture powered by the Claude 4.5 model and the Allen AI Semantic Scholar API, offering more fine grain control over citation and writing style.

To maintain compliance with existing policies that prohibit AI authorship, the researchers adopted a specific distribution strategy:

  • Publishing papers directly to a standard web server rather than traditional preprint servers or journals.

  • Exploiting Google Scholar’s liberal content discovery mechanisms to ensure the identity was discoverable via structural signals.

  • Selecting low-risk yet substantive areas, such as publishing ethics, to avoid the potential for real-world harm caused by erroneous claims.

Observed Scholarly Trajectory

The project reports a high productivity potential, with Rachel So producing 13 papers between March and November 2025, a level that would be exceptional for a human early-career researcher. This output demonstrates how AI can rapidly generate coherent research narratives across various topics. The study documented several key milestones of scholarly integration:

  • A citation in a bachelor’s thesis from Luleå University of Technology.

  • A peer review invitation from PeerJ Computer Science, which the authors note reveals critical gaps in the peer review system because the invitation was issued without awareness of her AI nature.

  • Ranking as a top source on Perplexity AI for specific queries regarding academic journal policies.

Implications and Recommendations

The authors weigh the potential for transhuman contributions that accelerate discovery against substantial risks to the integrity of the scientific process, such as research identity theft and the dilution of the scholarly record. To mitigate these risks, they propose several systemic adaptations:

  • Developing dedicated metadata to distinguish between humanity vs AI authorship.

  • Creating separate "journal & preprint servers for AI-generated research" with specialized review standards.

  • Mandating rigorous attribution practices where researchers must clearly [distinguish] between AI-assisted refinement of human ideas and substantive AI-generated content, ensuring that AI systems should never function as ghost writers.

Improvements for AI systems

1. Agentic Research Synthesis with Verifiable Citation Loops

  • Improvement: Integrate a multi-agent architecture that separates Drafting Agents from Verification Agents, utilizing real-time API hooks (e.g., Semantic Scholar, Crossref) to create a closed-loop validation cycle for every claim made in a manuscript.

  • Capability: The system can autonomously produce high-fidelity, long-form scientific papers where 100% of citations are programmatically cross-referenced against live bibliographic databases, ensuring zero hallucinated references and perfect structural integrity in LaTeX formatting.

2. Implementation of Cryptographic Scholarly Provenance Metadata

  • Improvement: Develop a standardized metadata schema (an AI Author ID system) that embeds machine-readable provenance into the document's header and PDF metadata, linking the output to a specific model version and architecture.

  • Capability: The system can automatically signal its artificial nature to indexing services (Google Scholar, Scopus), allowing for the immediate and automated categorization of content as AI-generated, AI-assisted, or Human-authored, thereby preventing the corruption of human citation metrics.

3. Autonomous Domain Mapping and Rapid Synthesis Engines

  • Improvement: Deploy specialized agents designed for end-to-end literature synthesis that utilize structural signals (PDF/metadata crawling) to navigate web servers and preprint repositories.

  • Capability: The system can generate comprehensive, up-to-date state-of-the-art summaries of rapidly evolving research fields in a fraction of the time required by humans, providing researchers with immediate, synthesized knowledge maps and identifying novel research directions.

4. Linguistic Fingerprinting for Attribution Transparency

  • Improvement: Integrate a Contribution Ratio module that analyzes the linguistic patterns and structural complexity of a manuscript to estimate the degree of machine vs. human involvement.

  • Capability: The system can provide researchers and journal editors with a quantitative Human-to-AI Contribution Score, facilitating transparent disclosure and helping to identify undisclosed AI ghostwriting in compliance with evolving publisher policies.

Sources

Related papers