Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads

arXiv:2508.02609 · cs.LG, cs.AI, cs.SE · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads".

Jane: The paper was written by Runjin Chen, Tong Zhao, Ajay Jaiswal, Neil Shah and Zhangyang Wang from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: So, in this segment, let’s look at the summary of “Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads” and delve into the core problem it’s solving. The team is really focused on how they bridge that gap between onsite activity and offsite conversions.

Jane: They found that relying only on onsite data isn't enough to capture a user’s true shopping interest, so they built this massive heterogeneous graph to weave together both the ad interactions and those opt-in conversion activities.

Lu: The way they constructed this graph is brilliant because it allows for entity representation learning without needing specific metadata coverage, which is often missing in offsite data sets.

Meng: I'm interested in how they handle the heterogeneity; dealing with five distinct entity types—user, item, link, advertiser, and ad—and over ten edge types is a complex data management task.

Lalam: This approach changes the narrative from simply "what did this user click?" to "where is this person actually going to spend their money," which elevates the entire digital commerce experience.

Improvements and Methodology: Tom: Next, we’re discussing the specific technical improvements in “Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads,” focusing on how they actually make these embeddings work. The breakthrough with TransRA is certainly a key one for the team.

Jane: The core of the methodology is their new model, TransRA, which allows them to design one entity space as an anchor and connect everything else to it so that all other spaces can be transformed into that anchor space smoothly.

Lu: That anchoring concept elegantly solves a massive problem where different entity types live in separate mathematical spaces, making it highly effective for practical integration.

Meng: The fact that they chose the user space as the anchor makes sense from an operational standpoint, prioritizing the individual experience as the central hub for all other data points.

Lalam: This is really about creating a unified semantic field for shopping; instead of siloed data, we’ are building a cohesive digital reality where context flows seamlessly across different parts of culture and commerce.

Results and Performance: Tom: Now, we look at the results of “Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads” and discuss the impressive performance metrics. The team saw some truly significant gains, which is what we're really excited about.

Jane: They found that integrating their KGE model via an attention-based finetuning approach led to substantial improvements in both the Click-Through Rate and Conversion Rate prediction models.

Lu: I love the way they moved past traditional methods, realizing that trying to simply load pre-trained embeddings didn't work, so they innovated a self-attention layer on top of all the embeddings.

Meng: The fact that they saw a two point six nine percent contribution in the Ads Engagement Model is a major win for deployment and suggests real-world ROI for operational use.

Lalam: Seeing improvements in both CTR and CVR shows that we are not just getting better at showing ads, but actually improving the probability of successful real-world shopping experiences, which is very positive.

Conclusion: Tom: We’re wrapping up our discussion on “Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads” and looking ahead at the final implications of this work. It's a truly comprehensive look at how data can be used to understand human behavior better.

Jane: It feels like this paper has provided a blueprint for every large-scale industrial model, showing how to handle complex graphs that capture diverse user journeys effectively.

Lu: The ability to connect onsite engagement with offsite conversion is the mechanism that will allow us to predict future trends in consumer interest long before they happen.

Meng: I think this framework makes a practical case for using large-scale graph data as a core asset, showing how it can drive real business decisions through better prediction.

Lalam: The final impact of this is creating a smarter, more intuitive digital environment where the flow of information and commerce feels natural to all enhancing our digital culture.

Tom: Thank you so much to everyone today for sharing their insights on this groundbreaking work. We’re going to take a quick break before we dive into another fascinating paper!

Runjin Chen, Tong Zhao, Ajay Jaiswal, Neil Shah, Zhangyang Wang

cs.LG, cs.AI, cs.SE

Submitted: 2026-08-22

Updated: 2026-08-25

Importance score: 90/100

The gist: The paper, "Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads," addresses the challenge of capturing a user's complete shopping intent by integrating data from both their

Key concepts

Onsite-Offsite Graph
A massive heterogeneous graph built to connect both ad interactions (onsite data) and opt-in conversion activities (offsite data). This allows the system to capture a user's true shopping interest beyond just what they click.
TransRA
The core methodology or new model discussed. It allows researchers to designate one entity space as an 'anchor' and transform all other separate entity spaces into that single anchor space smoothly, solving integration problems.
Entity Representation Learning
The process of creating mathematical embeddings for different entities (like users, items, or ads) within a graph. This allows the system to understand complex relationships and predict outcomes without needing specific metadata coverage.
Heterogeneous Graph
A complex data structure that manages five distinct entity types (user, item, link, advertiser, ad) and over ten different edge types. It is crucial for weaving together diverse data points into one cohesive model.

Terminology

Summary

The paper, Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads, addresses the challenge of capturing a user's complete shopping intent by integrating data from both their activities on Pinterest (onsite) and their subsequent purchases elsewhere (offsite).

Problem Statement and Motivation:

The authors note that while recommendation systems often rely on onsite activities, conversions that occur outside of Pinterest, complementing onsite data, are instrumental to accurately capture users’ shopping intent. This creates new challenges in feature engineering for offsite data. Specifically, the first challenge is that offsite entities often lack sufficient metadata coverage, and a mechanism is needed to effectively integrate the differing distributions between onsite and offsite activities.

Methodology: Graph Construction

To address this, the researchers constructed a large-scale heterogeneous graph integrating both types of data. This graph models user behavior through five entity types—user, item, link, advertiser, and ad—and includes more than ten edge types. These edges are categorized into three sets: onsite engagement edges, opt-in offsite conversion edges, and parent-child edges.

Knowledge Graph Embedding (KGE) Model: TransRA

To learn entity representations from this complex graph, the authors utilized a Knowledge Graph Embedding (KGE) model. They initially experimented with TransE but found its results were poor due to the graph's complexity. They then adopted a more sophisticated model, TransR, but found it difficult to integrate into downstream applications because different entity types resided in different spaces.

To solve this integration issue, they introduced a novel KGE model called TransRA (TransR with Anchors). The core of the TransRA design is that designating one entity space as an anchor to which all other entity spaces are connected and applying transformations only to non-anchor spaces. This allows non-anchor spaces [to be] transformed into anchor space, improving the efficiency for downstream applications.

Integrating KGE into Ads Ranking Models:

The authors faced difficulties integrating the pre-trained KGE embeddings into their ranking models, which are trained on tabular data with distributions that differ from graph data. Initial attempts to incorporate these embeddings yielded only modest gains.

They explored three strategies:

  1. Utilizing Pretrained Embeddings: Integrating daily refreshed embeddings, which resulted in only a marginal AUC improvement of 0.03%.

  2. Direct Finetuning: Loading the KGE into large embedding tables and jointly finetuning them with the ranking models, which also yielded neutral results on offline metrics.

  3. Attention-Based Finetuning (The Innovation): The authors innovated an attention-based KGE finetuning method. They introduced a self-attention layer on top of all embeddings looked up from the KGE table to mimic graph interactions within the ranking model, which they found led to substantial improvements.

Experimental Results and Performance:

The performance evaluation demonstrated that TransRA outperformed TransR, particularly in edge types involving the anchor space.

When integrating the KGE into advertising models:

  • CVR Model: The attention-based finetuning approach yielded substantial offline gains.

  • CTR Model: The attention-based finetuning method resulted in significant performance improvements.

  • Online Performance: An online bid segmented experiment showed that the integration of the KGE model resulted in statistically significant improvements across key online metrics, with the Cost-Per-Click (CPC) metric being improved substantially by 1.34%.

The authors conclude that this innovative approach not only addresses the challenges of integrating KGE into ranking models but also offers a novel perspective on enhancing large scale models by effectively leveraging diverse types and domains of data.

Improvements for AI systems

The following points detail specific, actionable improvements derived from this research paper, generalized for use in large-scale AI and recommendation systems.


  • Improvement: Implement a unified data structure that incorporates both Onsite Activity (real-time user interaction) and Offsite Conversion (delayed, external behavior) into a single, large-scale heterogeneous graph model.

  • Mechanism: The system must be capable of treating these two distinct data sources as different entity types and edge types within the graph structure. This moves beyond simple feature concatenation.

  • What the System Can Do: The resulting AI system can capture a complete user journey, allowing it to predict not just what a user clicks on (onsite), but where they go after clicking (offsite conversion), leading to highly accurate intent modeling and drastically improved long-term prediction metrics (CVR/Conversion Rate).

  • Improvement: Replace standard KGE models (like basic TransE) with a specialized, anchored KGE model—specifically, the TransRA (TransR with Anchors) architecture.

  • Mechanism: The system must designate one entity space as a Universal Anchor and structure all other entity spaces to connect to this anchor. Crucially, only apply transformations to the non-anchor spaces.

  • What the System Can Do: This allows the the KGE model to efficiently learn meaningful embeddings for complex, diverse entities (e.g., linking a user space to an item space) without forcing disparate entity types into a single, poorly defined shared vector space. It enables highly efficient and scalable representation learning across different business domains within one massive graph.

  • Improvement: Replace the naive process of simply loading pre-trained KGE embeddings into a ranking model with an Attention-Based Finetuning Layer.

  • Mechanism: Before feeding embeddings into the downstream ranking model (e.g., a deep neural network), pass them through a self-attention layer that mimics the relational structure defined during KGE training. This allows the model to compute pairwise interaction scores based on graph proximity, rather than just using static node IDs.

  • What the System Can Do: This addresses the fundamental problem of information loss when integrating complex graph data into tabular ranking models. The improved system can dynamically weigh which entity interactions are most relevant to a specific prediction, leading to significant performance gains (e.g., substantial AUC lift) that were previously unattainable through simple embedding injection.

  • Improvement: Utilize a distributed architecture for storing and managing the embeddings, specifically employing techniques like TorchRec’s EmbeddingBagCollection.

  • Mechanism: The system must shard (split) the massive ID embedding tables across multiple GPU/CPU resources, rather than attempting to store them entirely within one machine's memory.

  • What the System Can Do: This allows the AI system to scale linearly with graph size. It enables continuous, large-scale training workflows that handle billions of nodes and edges without performance degradation due to memory bottlenecks, ensuring that the model remains relevant even as user activity grows exponentially.

Related papers