Computer Science Conferences Should Require Nonrepudiable Experimental Results
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Computer Science Conferences Should Require Nonrepudiable Experimental Results".
Jane: The paper was written by Mamadou K. KEITA and Christopher Homan from Rochester Institute of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: The authors, Mamadou K. Keita and Christopher Homan, are proposing a huge change to the way research is conducted at conferences like NeurIPS or ICML.
Jane: They're arguing that we need a system where the numbers in a paper—the final results—are linked to an actual executed computation in a way that cannot be denied later by the authors.
Lu: That concept of "nonrepudiation" is really powerful; it means the author can't claim they didn't run the code or that they cheated.
Meng: It addresses the practical reality that if we don’t trust the input, we can’t trust any output, no matter how plausible it looks.
Lalam: The title itself forces us to ask: what kind of evidence do we accept in science today?
Tom: It makes you wonder about the culture of "publish-or-perish," doesn's it? Because if results are taken on faith, the pressure is immense.
Jane: And Keita and Homan seem to be saying that this lack of verifiable trust is a systemic failure, not just a few bad actors.
Lu: They aren't just blaming researchers; they are pointing out the structural weaknesses in how we verify claims.
Meng: If I were building a pipeline for this, I’d need to know precisely what level of proof was required before I even started writing code.
Lalam: The idea suggests that the integrity of our collective knowledge depends on a verifiable chain that goes from the initial hypothesis all the way through to the final reported metric.
Tom: So, we' are moving past just trusting the authors and towards some kind of cryptographic assurance for results.
Jane: Which brings us to how they think this problem actually looks when they break down what’s currently in place.
Summary of the Problem: Tom: In this section, we look at what Keita and Homan say is the "Verification Gap," which is essentially the gap between what we ask for and what we can actually check.
Jane: They found that all current verification tools share a common weakness: they are voluntary, self-reported, or only check functionality.
Lu: Even when researchers use documentation checklists like REFORMS, it doesn' to verify execution; it just confirms the checklist was filled out honestly.
Meng: The practical issue is that artifact evaluation checks if code runs at all, but not if the specific run produces the numbers cited in the paper.
Lalam: It’s like checking if a tool can be used to mix ingredients, rather than checking if it actually produced a cake with those exact properties.
Tom: And they bring up logging tools like Weights and Biases, which are useful for internal management but author-controlled, so the logs only show what the author *wants* us to see.
Jane: That’s a crucial distinction: if we rely on self-reported logs, we are relying on trust in an existing system rather than proof of actual event.
Lu: They also point out that pre-registration commits to a plan, but it doesn't bind the actual final output to that initial plan.
Meng: So, if I see code running successfully during artifact evaluation, I know the code *can* run, but not whether it executed the specific steps required to produce those reported accuracy scores.
Lalam: This means our current methods are mostly about checking intentions or abilities rather than establishing an undeniable historical record.
Tom: It’s a massive distinction between trusting a plan and trusting the results of an experiment, isn't it?
Jane: And knowing that these shortcomings exist, we can now look at the specific mechanism they propose to bridge this gap.
Improvements and Solutions: Tom: The next section introduces "Experiment Nonrepudiation," which is their formal definition of the problem class they want to solve.
Jane: They want a tamper-evident record that binds the reported metrics—the numbers in Table one or Table two—to a specific executed computation that cannot be altered or denied by the independent parties.
Lu: To achieve this, we need properties like "Data Blindness," meaning no observer can see the raw dataset, only its size and structure.
Meng: And crucially, the concept of "Execution-binding" means that if a paper claims GPU training on a large dataset, the system must prove it saw actual hardware activity consistent with that claim.
Lalam: This is where the cultural shift happens; we are moving from trusting paper narratives to trusting verifiable computational evidence.
Tom: The proposed solution is K-Veritas, and it’s a testbed implementation in Go that actually demonstrates how this could work at a basic level.
Jane: It wraps existing commands, captures all the necessary outputs, and then seals the entire session into a single cryptographic digest.
Lu: I find the idea of "Author-key separation" particularly interesting; keeping the private signing key outside the author is essential to ensure nobody can forge their own results.
Meng: From an engineering view, K-Veritas handles this by hashing everything—the source code, the configuration, and even tracking CPU time—before sending one single digest to a remote attestation service.
Lalam: The fact that the system generates a signed PDF report based on this record ensures that the final result is permanently anchored in a verifiable historical log.
Tom: It's not just about stopping small lies; it also addresses how we handle potential OS tampering or superficial fakes of training loops, which is huge.
Jane: But even though K-Veritas solves many problems, it doesn't solve everything, and that leads us to the big picture.
Conclusion and Wrap-Up: Tom: So, we’ve covered how the current system fails to verify results, what "nonrepudiation" is supposed to look like technically, and how K-Veritas provides a proof of concept.
Jane: We're looking at a future where verifying experimental integrity becomes a standard part of the academic workflow.
Lu: I’m excited about the potential for this to build trust across different fields that use computational pipelines, not just AI research.
Meng: The practicality is that by adopting this verifiable system, we reduce the risk of wasted follow-up research caused by unreliable published numbers.
Lalam: This paper on Computer Science Conferences Should Require Nonrepudiable Experimental Results ultimately suggests a more rigorous standard for how we validate knowledge in our culture.
Tom: It’s a call for community consensus to move towards an independent, non-profit organization to oversee this standard, right?
Jane: And while they acknowledge that some complex hardware attacks remain a challenge, the effort is significant.
Lu: The fact that they are being honest about what their current software tool can't prevent shows a high level of rigor.
Meng: It’s about making the cost of fabrication higher than the cost of doing real work, and that's a strong argument for me.
Lalam: By enforcing this standard, we are building a culture where evidence is not just accepted, but where we know it is verifiable truth.
Tom: So, as we wrap up our discussion on Computer Science Conferences Should Require Nonrepudiable Experimental Results, I think the message is clear: trust in science must be built on undeniable evidence.
Jane: It's a fascinating and important conversation to end with that nonrepudiation is becoming a central requirement.
Rochester Institute of Technology
cs.CR
Submitted: 2026-05-09
Updated: 2026-09-02
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: The paper, "Computer Science Conferences Should Require Nonrepudiable Experimental Results," addresses a critical failure point in modern machine learning research: the lack of verifiable trust in
Key concepts
- Nonrepudiation
- This concept means a system where an author cannot deny having run the code or claiming they cheated. It requires linking final results to an actual, undeniable executed computation that can be verified by independent parties.
- Verification Gap
- This refers to the gap between what researchers ask for and what current verification tools can actually check. Current methods often rely on self-reporting or only confirm a checklist was filled out, rather than proving the actual execution of results.
- K-Veritas
- This is a testbed implementation designed to solve the verification gap. It captures all necessary outputs from a computation, hashes everything—including source code and CPU time—and seals the entire session into a single cryptographic digest.
Terminology
Summary
The paper, Computer Science Conferences Should Require Nonrepudiable Experimental Results,
addresses a critical failure point in modern machine learning research: the lack of verifiable trust in reported experimental outcomes. The authors argue that existing reproducibility mechanisms are insufficient because they fail to cryptographically bind published results to the specific computational execution that generated them. By mandating nonrepudiation, the paper proposes a robust protocol designed to guarantee that reported numbers are tied irrevocably to an auditable, tamper-evident record of computation, thereby elevating scientific rigor and accountability across major conferences.
Comparison with Existing Approaches
The proposed nonrepudiation mechanism significantly surpasses current voluntary or partial reproducibility efforts. As detailed in the comparison table, nonrepudiation is the only method that achieves universal compliance and full assurance across all measured properties. Unlike existing tools such as W&B/MLflow or MLRC, which may only provide partial guarantees, nonrepudiation offers a comprehensive solution that is Tamper-evident
and ensures that the process can be verified during pre-review stages. Furthermore, the protocol mandates that Author cannot modify record,
providing an immutable chain of custody for results.
Core Properties of Nonrepudiation
The system is designed around several critical properties that establish trust and verifiability in scientific submissions. These include:
-
Binding Results: The protocol
Binds result to run,
ensuring that the reported numerical outcome cannot be detached from the specific execution environment and parameters used. -
Immutability: It is
Tamper-evident,
meaning any attempt to alter the recorded artifacts or results will be immediately detectable. -
Accessibility: The system requires
No data access required
for initial verification, streamlining the review process. -
Scope: Crucially, it provides assurance that is
Universal (all submissions),
making compliance mandatory and comprehensive across all submitted work.
Operational Requirements and Limitations
While highly robust, the authors acknowledge five significant limitations that must be addressed for successful production deployment at conference scale. These limitations highlight the operational complexity required to implement such a system:
-
Poor Experiment Design: Nonrepudiation only binds reported numbers to an actual execution; it
does not verify that the execution was well-designed.
A poorly controlled experiment, even with verified numbers, remains a poor scientific endeavor. -
Attestation Service Security: The
tamper-evidence guarantee depends on the security of the attestation service and its signing key.
Compromise of this service could lead to fabricated attestations, necessitatingfederation across multiple independent attesters
and rigorous key management. -
Author Compliance: The protocol is only effective when compliance is mandatory. If an author refuses to use a compliant implementation,
no attestation is generated,
limiting the system's reach without institutional mandate. -
Adversarial Attacks: A software-only observer cannot defeat all adversaries; specifically,
a software-only observer does not defeat OS-level or hardware-level adversaries.
The path toward high assurance requires considering hardware-backed attestation. -
Infrastructure Scale: Production deployment requires substantial institutional infrastructure, including
persistent session storage, rate limiting, key rotation, and auditing,
which are operational challenges rather than protocol flaws.
Improvements for AI systems
The current body of work establishes that reproducibility is not a single technical fix but a multi-layered scientific process requiring mandatory infrastructure and behavioral changes. The proposed improvements focus on integrating nonrepudiation and security principles directly into the AI research workflow to create a verifiable, auditable, and trustworthy research ecosystem.
Here are the specific improvements for an AI system:
Improvement: Develop a mandatory, blockchain-backed ledger that serves as the single source of truth for all components of an AI experiment. This system must enforce cryptographic binding between five distinct artifact types at the point of submission:
-
Data Snapshot: Immutable hash and provenance chain (source, cleaning steps, version).
-
Code Repository: Version-controlled code (e.g., Git commit hash) with mandatory dependency pinning (
condaenvironment specification). -
Hyperparameters/Config: Structured JSON file detailing every tunable parameter used.
-
Execution Environment: Container image hash (e.g., Docker/Singularity) ensuring OS and library versions are fixed and verifiable.
-
Result Payload: The final output metrics, which must be cryptographically bound to the execution environment via a nonrepudiation signature (see point 3).
What the Improved System Can Do:
-
Guaranteed Traceability: It eliminates ambiguity regarding which specific combination of data, code, and environment produced a reported result. A user can query the IRAL and retrieve the exact computational context used for any published claim.
-
Automated Sanity Checks: The system can flag submissions where the input dependencies (e.g., a library version) are known to be susceptible to recent CVEs or major breaking changes, prompting the author for justification or mitigation.
Abstract
This position paper argues that computer science conferences should require tamper-evident, nonrepudiable attestations of experimental results. We name the underlying problem experiment nonrepudiation: a compliant protocol must bind the numbers in a paper to an actual executed computation in a way the author cannot later alter or deny. The current system relies on self-reported checklists, optional code sharing, and author-controlled logging. None of these mechanisms answer the question a reviewer cannot check: did the code the paper describes produce the numbers the paper reports? We define the problem formally, state the security properties any compliant protocol must satisfy, and describe a threat model that includes attacks current approaches do not prevent. We frame the mechanism that provides this property as a proof-of-compute layer, a general requirement that computed results be admitted as evidence only when accompanied by a verifiable record of the computation that produced them. To show that the problem is solvable, we built K-Veritas, a reference implementation in Go that produces signed reports without accessing training data. K-Veritas is a testbed, not a finished answer. We call on conferences and the community to treat nonrepudiation as a major requirement and to help build an open, independent standard for it.
Sources
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs