ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion".
Jane: The paper was written by the authors from Institute of Software, Chinese Academy of Sciences and University of Chinese Academy of Sciences and CITIC Securities and School of Artificial Intelligence, Beijing Normal University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 2: Tom: Last time, we established that "ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion" fundamentally links sight and language to structure knowledge. Now, let’s look closer at the summary provided by the authors—what does this system actually do when it processes data?
Jane: If we focus on the core mechanism, the paper describes how ViSR-KGC takes multiple inputs simultaneously: an image, a chunk of text, and an existing knowledge graph. The model doesn't process them in isolation; it forces them to interact during the reasoning step.
Lu: That interaction is what allows for deep inference. Instead of just running object recognition on the photo, or simple entity extraction on the text, the model uses one input to constrain and improve its understanding of the other two. It’s a feedback loop of interpretation.
Meng: To elaborate on that constraint: imagine an image shows a piece of machinery, and the text describes its function. The model doesn't just say "It's machinery" or "Its function is X." It uses the *visual evidence* to confirm if the described function is plausible for what it sees, and vice versa.
Lalam: And this means that even if the text is vague or slightly misleading, the visual context can act as a powerful anchor, pulling the entire interpretation back towards something more accurate and grounded in reality.
Tom: So, we're moving past simple data linkage; it’s about deep integration of reasoning across multiple data types at once.
Jane: Precisely. It’s about constructing a reasoning path that is informed by all available perspectives—the visual relationships, the textual definitions, and the established structure of the graph itself. It helps fill in those gaps where knowledge was never explicitly written down or photographed clearly enough before.
Lu: I think thinking about it as a process of *hypothesis generation* is helpful here. The model isn't just retrieving facts; it's proposing new, highly probable relationships based on the combined weight of all its inputs.
Meng: And this level of integrated reasoning means that the system is inherently more robust to ambiguity than older methods that often broke down when data was even slightly contradictory or incomplete in one area.
Lalam: It’s a powerful demonstration that human understanding rarely comes from a single source, but from the synthesis of many different kinds of sensory and conceptual inputs.
Tom: This deep synergy between vision, language, and structure is truly remarkable. But what does this capability translate into when we consider the real-world potential?
Jane: It suggests that we are moving toward systems that don't just process data points, but understand the underlying narrative or history connecting those points.
Paper discussion segment 3: Tom: We’ve established how "ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion" combines inputs to reason deeply. Now, let's focus specifically on the *improvements* this architecture suggests for making AI smarter in the real world. What are the practical leaps?
Jane: The biggest improvement, as I see it, is that it solves what they call "contextual gaps." Most current systems fail when information is presented messily or partially. ViSR-KGC’s breakthrough improvement is its ability to use the visual arrangement to *predict* what crucial missing link must exist for the whole picture to make sense.
Lu: To build on that, think of it like a detective looking at a crime scene photo—they don't just list every object they see. They use the spatial relationship between those objects, maybe an angle or a shadow, to infer that something *must* have happened there, even if no explicit evidence remains. The model is learning to build the missing narrative structure itself.
Meng: From an engineering standpoint, this leap suggests a massive improvement in data ingestion pipelines. Instead of needing perfectly labeled nodes and edges—which is impossible for real-world datasets—the system can operate on low-fidelity inputs: sketches, overheard notes, or grainy photos. It’s moving us from requiring *perfect* data to accepting *imperfect* reality.
Lalam: And that democratization of input is huge because it expands who can use the tool. Right now, advanced AI requires highly curated academic datasets to function optimally. The suggested improvements mean that a person with a specialized skill—say, an archaeologist pointing out an unusual tool—can use the system, and the AI can understand the context without needing years of pre-training on that specific culture’s data.
Jane: So, in simple terms, it makes AI less of a bookkeeper—someone who just checks off facts—and more like a highly intuitive research assistant who anticipates your next question based on what you’ve already shown them. The visual evidence guides the entire line of questioning.
Lu: It transforms
Paper discussion segment 3: Tom: To recap our journey through ViSR-KGC, we’ve established that its fundamental strength is moving beyond simple data linkage to achieve deep, integrated reasoning using visual and textual inputs simultaneously. Today, let's focus purely on the implications—what specific *improvements* does this architecture suggest for making AI smarter in the real world?
Jane: When I think about the practical improvements suggested by the authors, it seems they are solving the problem of "contextual gaps." Most current systems fail when information is presented messily or partially. ViSR-KGC’s improvement is that it doesn't just look at individual pieces of evidence; it uses the visual arrangement to *predict* what crucial missing link must exist for the whole picture to make sense.
Lu: Exactly. Think of it like a detective looking at a crime scene photo—they don't just list every object they see. They use the spatial relationship between those objects, maybe an angle or a shadow, to infer that something *must* have happened there, even if no explicit evidence remains. The model is learning to build the missing narrative structure itself.
Meng: From an implementation standpoint, this leap suggests a massive improvement in data ingestion pipelines. Instead of needing perfectly labeled nodes and edges—which is impossible for real-world datasets—the system can operate on low-fidelity inputs: sketches, overheard notes, or grainy photos. It’s moving us from requiring *perfect* data to accepting *imperfect* reality.
Lalam: And that democratization of input is huge. Right now, advanced AI requires highly curated academic datasets to function optimally. The suggested improvements mean that a person with a specialized skill—say, an archaeologist pointing out an unusual tool—can use the system, and the AI can understand the context without needing years of pre-training on that specific culture’s data.
Jane: So, in simple terms, it makes AI less of a bookkeeper—someone who just checks off facts—and more like a highly intuitive research assistant who anticipates your next question based on what you’ve already shown them. The visual evidence guides the entire line of questioning.
Lu: It transforms the process from "Here is all the data; find an answer" to "Look at this picture; what questions should we ask to understand this object's full history?" That shift in agency is where the real power lies.
Meng: And that guided questioning capability means we can build diagnostic tools for fields like medicine or engineering, where a partial scan or a rough prototype needs immediate, intelligent analysis rather than just a database search.
Tom: This ability to guide discovery is transformative. It suggests that the next frontier isn't just making AI bigger, but making it much more context-aware and intuitive in its reasoning process. Speaking of intuition, this brings us to the boundary where AI starts generating content itself—what happens when perception merges with creation?
Conclusion: Tom: So, as we wrap up our deep dive today, what really strikes me is that the shift isn't just about better data processing; it’s about fundamentally changing how machines approach understanding complex reality.
Jane: Exactly. We’ve moved past the idea of AI merely finding answers and toward building systems that genuinely reason, synthesizing disparate pieces of evidence across vision, text, and relationships simultaneously.
Lu: For me, the deepest implication is how this framework could reshape scientific collaboration—allowing researchers to formulate hypotheses based on visual data that they might not have been able to articulate in purely textual terms before.
Meng: From a broader industrial standpoint, I see the immediate value in automating domain expertise. This capability means that complex knowledge previously locked away behind academic journals can now be accessed and utilized through these integrated models.
Lalam: And I keep coming back to the idea of accessibility; this technology has the power to democratize understanding, giving people from all backgrounds a way to interact with deep, structured knowledge they might not otherwise have access to.
Tom: Lalam’s point about democratization is profound. It suggests that this isn't just an academic tool, but a potential catalyst for cultural and educational shifts across the globe.
Jane: It really boils down to recognizing that human intelligence rarely comes from one single source, but from the constant cross-referencing of many types of inputs—and that’s what *ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion* achieves so brilliantly.
Tom: It has given us a powerful blueprint for how future AI systems need to operate if they are ever going to be truly useful partners in solving the world's toughest problems.
Jane: It’s been an incredibly informative and insightful discussion, everyone; thank you all for walking us through the implications of this groundbreaking work today.
Tom: Alright listeners, that wraps up our deep dive on this paper for now; next week, we're shifting gears entirely and looking at how AI is tackling personalized synthetic media generation—you won't want to miss it!
Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · CITIC Securities · School of Artificial Intelligence, Beijing Normal University
cs.AI
Submitted: 2026-08-06
Updated: 2026-09-09
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 95/100
The gist: The paper, published at ACM MM '26 by Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, and Hongan Wang, addresses the task of multimodal knowledge graph completion (MMKGC).
Key concepts
- Multimodal Knowledge Graph Completion
- This process uses multiple types of input—like images and text—to fill in missing information (gaps) within a structured knowledge graph. It moves beyond simple fact retrieval to synthesize complex relationships.
- Visual Subgraph Reasoning
- The system uses visual evidence and spatial relationships, like those seen in a photo, to infer missing facts or connections. Instead of just listing objects, it builds the missing narrative structure based on what is seen.
- Deep Integration of Reasoning
- This refers to the model's ability to force an image, a chunk of text, and a knowledge graph to interact simultaneously. One input helps constrain and improve the understanding derived from the other two inputs.
Terminology
Summary
The paper, published at ACM MM '26 by Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, and Hongan Wang, addresses the task of multimodal knowledge graph completion (MMKGC). The authors note that "Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images."
The paper identifies two key limitations of existing approaches:
-
Traditional embedding-based methods:
Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited.
Such methods fail particularly on few-shot or zero-shot relation types, as illustrated by theportrayer
relation example in Figure 1. -
LLM-based reasoning methods: These
typically linearize graph structures into textual prompts, which obscures structural topology and neglects vital visual information.
Meanwhile, the authors observe that vision-language models (VLMs) excel at multimodal reasoning, they cannot natively interpret structured graph topology, particularly when it comes to knowledge graphs where nodes and edges carry complex semantics.
This motivates the central research question: in the VLM era, can knowledge graph completion, particularly for multimodal knowledge graphs, directly benefit from visual representations of graph structures?
The paper identifies three key challenges: (1) subgraph extraction — VLMs can only process a single image, so a compact, query-relevant subgraph must be extracted; (2) subgraph visualization — the subgraph must be rendered in a layout that enables VLMs to effectively identify triples, understand their correlations, and capture associated textual/visual semantics
; and (3) prompt integration — the prompt must "jointly incorporate local structural information from the subgraph and global features from representation learning, while harmonizing the interplay between the large structural subgraph image and small embedded entity images."
The paper makes three main contributions:
-
First formulation of MMKGC as visual subgraph reasoning for VLMs: "This is the first work to formulate MMKGC as a visual subgraph reasoning problem for VLMs, through effectively integrating three complementary capabilities: multimodal representation learning to capture global semantics, VLM-based reasoning to extract local relational evidence, and VLMs' internal memory to provide supplementary commonsense knowledge."
-
A systematic pipeline that "extracts a query-aware subgraph based on graph structure as well as node semantics, and converts it into a visually interpretable image, facilitating VLMs to jointly reason over graph topology and multimodal entity information."
-
Extensive experiments on two real-world MMKG datasets demonstrating
the coordinated synergy between relational reasoning and commonsense knowledge enabled by VLMs.
An MMKG is defined as a quadruple G = (E, R, T, M)
where E is entities, R is relations, T ⊆ E × R × E is the set of factual triples, and M = M e denotes multimodal attributes for each entity. The task is to predict the missing entity: tˆ = argmax F (h, r, t ′)
for tail prediction, with the symmetric formulation for head prediction.
The framework consists of four stages:
1. Multimodal Entity Representation. The authors adopt IMF as the multimodal encoder, which provides effective joint representations across textual, visual, and structural modalities, achieving near-best prediction performance among embedding-based methods.
After training, each entity is associated with four embeddings: a textual embedding x t, a visual embedding x v, a structural embedding x s, and a fused multimodal embedding x m.
They note this component can in principle be replaced by other representation learning models.
2. Query-aware Subgraph Extraction. Since directly reasoning over the entire knowledge graph is infeasible for VLMs,
the method extracts a compact subgraph containing three categories of candidate edges:
-
Relation-consistent edges (C rel):
edges whose relation type r i matches the query relation r, while the head entity differs from the query entity
— providingadditional instances of the relational pattern.
-
One-hop neighbor edges (C 1-hop):
edges directly connected to the query entity h
— offeringthe most immediate structural context.
-
Two-hop neighbor edges (C 2-hop):
edges connected to the one-hop neighbors of h, but excluding those that still contain the query entity h
— capturingindirect structural relations and potential multi-hop reasoning paths.
Each candidate edge is scored via: score(h i, r i, t i) = λ sim(r, r i) + (1 − λ) max[sim(h, h i), sim(h, t i)],
where entity similarity is computed as the average cosine similarity across the four embedding views. A structural bonus is added for edges sharing an element with the query: score∗(h i, r i, t i) = score(h i, r i, t i) + γ · δ((h i, r i, t i) ∈ C 1-hop ∪ C rel).
3. Subgraph Visualization. The extracted subgraph is rendered as an image via graph rendering operation R(·)
using Graphviz. Based on empirical comparison of layout algorithms (dot, circo, twopi, neato, fdp, sfdp), the authors select dot as the default: the dot layout arranges nodes in a hierarchical manner and reduces edge crossings, making the resulting graph structure more readable and interpretable.
4. VLM Reasoning. The prompt combines textual input (the incomplete query triple, a textual serialization of the subgraph's key triples, the candidate entity set computed by representation learning, and task instructions) with visual input (the rendered subgraph image I sub and entity images I e). The prediction follows: eˆ = arg max P(e i T prompt, I sub, I e e∈V sub).
The candidate list serves as guidance rather than a hard constraint.
The VLM backbone is Qwen3VL.
Experiments are performed on FB15K-237 and DB15K, two widely used MMKG datasets. The test sets undergo strict preprocessing... retaining only triples with a unique correct answer for each query
to remove ambiguity. FB15K-237 has 14,541 nodes, 237 relation types, and 272,115/17,535/6,102 train/val/test triples; DB15K has 14,777 nodes, 279 relations, and 56,881/9,903/8,355 triples. Evaluation uses Hit@1 and Hit@3, reported separately for head and tail predictions. Hyperparameters: k max = 15 edges, k rel = 10, k nei = 10, λ = 0.5, γ = 0.1, on an NVIDIA RTX 4090.
ViSR-KGC achieves the best performance across all metrics and datasets, e.g., FB15K-237 tail Hit@1 = 0.8027 and Hit@3 = 0.8836; DB15K tail Hit@1 = 0.6715 and Hit@3 = 0.7802. The paper draws three key observations:
-
Traditional embedding-based methods struggle with complex reasoning. The gap confirms that
pure representation learning, which relies mainly on graph structure and embedding spaces, struggles to capture implicit relational patterns.
Notably, "ViSR-KGC leverages commonsense knowledge beyond candidate sets obtained from the representation learning. Around 9.96% and 11.24% of correct predicted entities on FB15K-237 and DB15K respectively, fall outside the candidate set.An example is given of correctly answering (?, creator, Mike Henry(voice actor)) with
This Is The Cleveland Show" by analyzing related entities (Family Guy, American Dad) plus internal knowledge of television animation production. On sparse relation types (fewer than quartile frequency, 7.6%/6.4% of edges), IMF achieves Hit@1 of 77.14%/37.5% versus ViSR-KGC's 85.16%/48.44%. -
LLM-based methods lack explicit graph structure perception. Even multimodal LLM methods
process graph structure implicitly through feature propagation
andremain below our ViSR-KGC.
The paper's case study shows that for query (University of Pittsburgh, affiliation, ?), an LLM predictsPennsylvania
(mistaking it for geographic association), whileViSR-KGC successfully identifies the correct entity by inspecting the visualized subgraph structure
— specifically revealing the pathUniversity of Pittsburgh → Pennsylvania → Pennsylvania State University → Association of American Universities,
whichimplicitly encodes the affiliation relationship.
-
Asymmetry between head and tail prediction.
All methods perform significantly better on tail entity prediction than on head entity prediction
becausehead entities are often more specific subjects (e.g. a person or organization), and tail entities are more general objects.
The ablations validate each component:
-
W/o relation-consistent edges (C rel) drops tail Hit@1 to 0.6232; w/o two-hop edges (C 2-hop) drops to 0.6598, confirming multi-type edges provide
focused evidence.
-
W/o subgraph entirely drops tail Hit@1 to 0.4545, showing subgraph structure is
critical.
Removing only subgraph text drops to 0.5075, while removing only the subgraph image drops to 0.6667 — text is more critical, but the image addsadditional structural cues.
-
Layout changes (twopi/sfdp) worsen accuracy, confirming
rendering the subgraph with an appropriate layout is crucial.
-
W/o candidate entities drops tail Hit@1 to 0.6533, showing the list
guide[s] the reasoning space of VLMs
and prevents hallucination. -
W/o entity images drops to 0.6566, showing entity images provide
complementary entity-level visual features.
-
Replacing the encoder with ConvE, ConvKB, or HGNN-IMA substantially lowers accuracy (best among them: tail Hit@1 = 0.6392), confirming
the quality of multimodal representations directly impacts the subgraph construction.
-
k max: Performance improves from k max=10 to 15, then Hit@3 slightly decreases at 18 because
excessive subgraphs may introduce noise and increase visual complexity.
Default k max=15. -
λ: Both metrics increase as λ rises from 0.3 to 0.5, then decrease beyond 0.5;
relation similarity is indeed beneficial, but over-emphasizing this measure would weaken the model's ability to distinguish entities.
Default λ=0.5.
The paper concludes that ViSR-KGC effectively harnesses VLMs' capabilities in multimodal understanding and associative reasoning grounded in internal memory
and integrates textual, visual, and structural information, supported by a carefully chosen multimodal encoder and the layout strategy.
The synergy enables the capture of three levels of semantic correlations for accurate link prediction: global and local evidence within the knowledge graph, as well as implicit commonsense knowledge residing in VLMs.
Future work directions include "more efficient subgraph extraction strategies to handle larger-scale multimodal knowledge graphs, integrating additional modalities such as audio or video to enrich multimodal semantics, and adapting the approach to other reasoning tasks on semantic graphs, as well as
adapting VLMs to MMKGC through reinforcement learning with verifiable rewards derived from link prediction correctness."
Improvements for AI systems
Based on ViSR-KGC, I would make the following concrete improvements to AI systems:
- Add explicit visual subgraph reasoning to VLMs. Instead of linearizing knowledge graph triples into text, render a query-relevant subgraph as an image and feed it directly to the VLM along with entity images. Use a hierarchical, low-edge-crossing layout (like Graphviz
dot).
Resulting capability: The system can perceive graph topology—paths, branching, and relation context—rather than only sequential text, enabling it to correctly answer multi-hop queries such as “University of Pittsburgh → Pennsylvania → Pennsylvania State University → Association of American Universities” that text-only LLM prompting gets wrong.
- Use query-aware subgraph extraction with heterogeneous edge types. Select three complementary edge sets: relation-consistent edges (same relation, different subjects), one-hop neighbor edges, and two-hop neighbor edges. Score each candidate edge by a weighted combination of relation similarity and entity embedding similarity across text, image, structure, and fused modalities, plus a structural bonus for edges sharing the query head.
Resulting capability: The system can reason effectively on sparse or few-shot relations, focusing on the most relevant evidence while avoiding irrelevant graph noise.
- Combine global embedding scores, local subgraph evidence, and VLM internal commonsense in one prompt. Use a multimodal representation model (e.g., IMF) to produce global entity candidates; render the subgraph image; provide the textual triple list; include the candidate set as soft guidance rather than a hard filter.
Resulting capability: The system can make correct predictions outside the embedding-based candidate set—about 10–11% of correct answers in the paper—by leveraging VLM commonsense knowledge (e.g., inferring “This Is The Cleveland Show” as creator of a voice actor from related TV shows), while the candidate list still reduces hallucination.
- Provide both subgraph image and subgraph text simultaneously. Do not rely on image alone or text alone. Keep the visual subgraph for structural cues and the textual serialization of triples for precise relational facts.
Resulting capability: The system remains accurate even when image resolution or graph rendering is imperfect; ablations show that losing subgraph text drops tail Hit@1 from 0.8027 to 0.5075, while losing only the image still degrades performance, confirming the need for both modalities.
- Use layout-aware graph rendering tuned for VLM perception. Choose a layout that minimizes edge crossings and hierarchical ambiguity (dot layout in the paper), and validate layout choice empirically.
Resulting capability: The system can reliably read graph structure from rendered images; switching to twopi or sfdp layouts measurably reduces accuracy, so the improved system treats visualization as part of the reasoning model, not as a fixed preprocessing step.
- Make the multimodal encoder replaceable and representation-quality-aware. Train an encoder that produces textual, visual, structural, and fused embeddings, then use these embeddings both for scoring and for subgraph candidate selection.
Resulting capability: The system can be upgraded as better multimodal encoders appear; higher-quality embeddings directly improve subgraph construction and downstream link prediction (replacing the encoder with weaker models drops tail Hit@1 from 0.6715 to 0.6392 or less).
- Handle head prediction and tail prediction asymmetrically. Because head entities are often specific subjects (people/organizations) while tails are general objects, generate direction-specific prompts, candidate scoring, and subgraph emphasis for head prediction.
Resulting capability: The system can close the gap between head and tail prediction accuracy, improving overall KG completion rather than optimizing only the easier tail direction.
- Incorporate scalable subgraph extraction and RL-based adaptation for larger graphs. Replace exhaustive candidate scoring with scalable sampling or hierarchical subgraph retrieval; optionally fine-tune the VLM with reinforcement learning using link-prediction correctness as a verifiable reward.
Resulting capability: The system can scale to larger multimodal knowledge graphs and continuously improve its reasoning from its own correct/incorrect predictions.
Overall, an AI system improved with these changes can perform multimodal knowledge graph completion by jointly reasoning over global embedding semantics, local visual graph structure, entity images, and VLM internal commonsense—yielding substantially higher Hit@1 and Hit@3 on benchmarks like FB15K-237 and DB15K, especially for sparse relations and multi-hop relational queries.
Abstract
Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images. Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited. Meanwhile, LLM-based reasoning methods typically linearize graph structures into textual prompts, which obscures structural topology and neglects vital visual information. While vision-language models (VLMs) excel at multimodal reasoning, they cannot natively interpret structured graph topology, particularly when it comes to knowledge graphs where nodes and edges carry complex semantics. To bridge this gap, we propose ViSR-KGC, a visual subgraph reasoning approach for KGC. It integrates three complementary capabilities to capture semantic correlations: identifying global topology dependencies via representation learning, analyzing local multimodal evidence using VLMs, and providing necessary commonsense knowledge inherent in pre-trained models. Based on learned multimodal embeddings, our framework first extracts a compact and query-aware subgraph from the MMKG. Then, this subgraph is transformed into a visually interpretable image using a layout strategy selected through empirical comparison. Finally, the visualized subgraph, entity images, textual descriptions, and candidate answers are combined into a unified prompt, enabling the VLM to infer the missing entity.
Sources
- Qwen3-VL Technical Report
- Relational Graph Attention Networks
- MLaGA: Multimodal Large Language and Graph Assistant
- ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion
- Knowledge Graph Reasoning with Self-supervised Reinforcement Learning
- Mario: Multimodal Graph Reasoning with Large Language Models
- KICGPT: Large Language Model with Knowledge in Context for Knowledge Graph Completion
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection