Learning with Local Search MCMC Layers

arXiv:2505.14240 · cs.LG · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Learning with Local Search MCMC Layers".

Jane: The paper was written by Germain Vivier-Ardisson, Mathieu Blondel and Axel Parmentier from Google DeepMind and ENPC.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary of the Approach: Tom: So, the core problem is that traditional methods for solving these hard problems, like finding a MAP solution, usually require an oracle—a perfect solver—which is often unavailable. This paper proposes a way around that by leveraging MCMC.

Jane: It essentially turns those iterative local search steps into a Markov chain process where we sample solutions from the combinatorial space, which allows us to treat the whole process as a single layer in our network.

Lu: The elegance lies in realizing that the neighborhood system—the set of available moves from being at one point to another—is exactly what defines how we should construct these proposal distributions for our MCMC sampler.

Meng: It’s critical to understand if this transformation actually reduces the computational burden. Are we just replacing one slow process with a faster, approximate one, or does it offer a genuine speedup?

Lalam: The impact here is that allows us to move away from the expectation of perfect optimization toward accepting good enough solutions, which is more reflective of reality in many dynamic situations.

Improvements and Mechanism: Tom: The authors achieved something quite impressive: they successfully made this entire combinatorial layer differentiable and showed it yields a stochastic gradient for a Fenchel-Young loss. That’s massive because gradients are what allow us to train the model!

Jane: The Fenchel-Young loss provides a principled way to define error when we can't guarantee an exact solution, allowing us to learn even if our approximate MCMC layer is imperfect.

Lu: I love the idea of Algorithm two the neighborhood mixture. By mixing multiple neighborhood systems, they ensure that even if one single movement type doesn's connected the entire space, a diverse set of movements can create a unified Markov chain that covers everything.

Meng: The experimental results in Section five point one are very encouraging; Table three shows a significant performance gain over perturbation methods, especially when we have tight time budgets for the layer's forward pass.

Lalam: This efficiency suggests we can tackle massive, real-world problem instances that were previously too large to handle without sacrificing quality of solution.

Conclusion and Future Outlook: Tom: We’ve seen how "Learning with Local Search MCMC Layers" moves us from the theoretical requirement of perfect solvers to a practical reality where we can use high-quality approximations.

Jane: It's comforting to know that even with inexact solvers, we can still achieve principled learning and that the authors have provided convergence analyses for these stochastic gradient algorithms.

Lu: I see this as a foundation for much more complex architectures, allowing us to build systems where decision-making is not just learned but is structurally grounded in real combinatorial constraints.

Meng: The work on initialization—specifically, the data-based initialization outperforming random starts—gives me confidence that when deploying these practical AI solutions, we have clear operational guidance for better performance.

Lalam: To wrap up, this framework helps us achieve a more nuanced form of intelligence in our systems, moving beyond simplistic decision-making to embrace the complexity of the world.

Tom: A fantastic discussion on "Learning with Local Search MCMC Layers." It's clear this is going to change how we approach complex optimization problems in AI.

Jane: Absolutely, it opens up a huge space for future development in structured prediction.

Lu: I'm excited to see what creative ways we can expand these mixed neighborhood systems into even more powerful architectures.

Meng: For the engineering side, having a reliable way to use local search heuristics is very practical for deployment at scale.

Lalam: I hope this innovation helps us build AI that respects the constraints of the real world, improving our global ability to solve problems together.

Conclusion: Tom: So, we've covered a lot today, but it really comes down to this: "Learning with Local Search MCMC Layers" gives us a genuinely principled way to solve hard combinatorial problems in AI.

Jane: It’s amazing how effectively the authors managed to bridge the gap between these complex local search heuristics and the beautiful mathematical structure of MCMC sampling, making it accessible for training.

Lu: I think we've really glimpsed something truly wild here, Jane; imagine building a system that can spontaneously explore vast solution spaces without getting stuck in those local optima traps we usually face.

Meng: And from an engineering standpoint, Lu’s point is spot-on because it means our AI can actually handle massive industrial problems at scale without needing a supercomputer running an exact solver all the time.

Lalam: I hope this technology allows us to build systems that don't just give answers, but that truly explore the breadth of possibilities, leading to more thoughtful and resilient solutions for every single user.

Tom: That’s a great way to put it, Lalam; we’re not just optimizing; we’re broadening the scope of what we can solve.

Jane: And by making this approach differentiable, as they showed in the paper, they made sure that the approximation actually fits seamlessly into our neural network training pipeline.

Meng: The practical impact is clear: where previous methods struggled with complexity or time constraints, this method provides a robust alternative that operates at a much faster pace.

Lu: It feels like we’ve unlocked a whole new family of possible architectures—one that respects the "rules" of the world while still learning from them.

Tom: I can't wait to see how these concepts evolve in future work, but it's been fascinating to break down this paper with all of you.

Jane: It’s definitely a conversation worth having again, Tom; we’ve learned that we have a whole new tool in our toolkit for the challenge.

Tom: We'll be back next time with another exciting discovery on the arXiv, so stay tuned for whatever comes next!

Germain Vivier-Ardisson, Mathieu Blondel, Axel Parmentier

Google DeepMind · ENPC

cs.LG

Submitted: 2026-08-24

Updated: 2026-08-25

Importance score: 90/100

The gist: This paper introduces a theoretically-principled framework for integrating NP-hard combinatorial optimization problems into neural networks by using local search heuristics as differentiable MCMC

Key concepts

MCMC Layers
The paper uses Markov Chain Monte Carlo (MCMC) to transform iterative local search steps into a single layer within the network. This approach allows the system to sample solutions from the combinatorial space, treating the entire process as one unified layer.
Fenchel-Young Loss
This loss function is used to define error when an exact solution is unattainable. It provides a principled way for the model to learn even if its approximate MCMC layer is imperfect, allowing for principled learning under approximation.
Stochastic Gradient
By making the entire combinatorial layer differentiable, this yields a stochastic gradient. This gradient is essential because it allows researchers to successfully train the AI model using this new architecture.
Neighborhood Mixture
This technique mixes multiple neighborhood systems together. It ensures that even if one movement type does not connect the whole space, a diverse set of movements can create a unified Markov chain.

Terminology

Summary

This paper introduces a theoretically-principled framework for integrating NP-hard combinatorial optimization problems into neural networks by using local search heuristics as differentiable MCMC layers. This matters because many real-world operations research problems are intractable for exact solvers, and existing inexact approaches often lack formal theoretical guarantees or fail during end-to-end training. By transforming neighborhood systems into proposal distributions, the authors enable efficient, scalable learning for complex structured prediction tasks.

The Core Problem

The researchers address the challenge of integrating combinatorial optimization layers of the form:

yb: θ ↦ argmax ⟨θ, y⟩ + φ(y)

where the output must satisfy complex constraints. Because these functions are piecewise-constant, they break the differentiable computational graph, preventing meaningful backpropagation. While some methods use regularization to create continuous relaxations, they often require an exact oracle to provide theoretical guarantees. However, in practice, many applications must rely on local search heuristics (like simulated annealing) which are inherently inexact and difficult to differentiate through directly.

How it works

The authors propose leveraging the link between local search heuristics and Markov chain Monte-Carlo (MCMC) methods. They transform problem-specific neighborhood systems into proposal distributions, effectively turning a local search oracle into a discrete MCMC sampler over the combinatorial set of feasible solutions. This allows for the construction of differentiable combinatorial layers that provide stochastic gradients. The framework handles two primary settings:

The unconditional setting:

Learning parameters θ from observations of feasible solutions.

The conditional setting:

Learning parameters W where θ = gW(x), mapping inputs to structured outputs.

To handle complex solvers, they introduce a method for mixing neighborhood systems. Instead of a naive aggregation that is computationally prohibitive, they propose an update rule that only requires computing the single individual ratio for the specific move type sampled. This allows the model to leverage a diversity of neighborhood systems while preserving the correct stationary distribution.

Theoretical Contributions and Loss Functions

The paper provides a rigorous mathematical foundation for this approach by deriving principled loss functions based on Fenchel-Young theory. Key theoretical results include:

Differentiability:

The proposed layer is differentiable, with its Jacobian given by the covariance matrix of the stationary distribution.

Convergence:

The authors provide a convergence analysis for stochastic gradient algorithms in both conditional and unconditional settings, proving that even with a single MCMC iteration (K=1), one can obtain an unbiased gradient estimator for a target-dependent Fenchel-Young loss.

Asymptotic Normality:

They prove that the empirical loss minimizer converges to the true parameter as sample sizes increase.

Empirical Validation

The effectiveness of the approach is demonstrated through large-scale experiments, most notably on a dynamic vehicle routing problem with time windows (DVRPTW). The results show:

Efficiency:

In low time-limit regimes (1–100 ms), the proposed method significantly outperforms the perturbation-based method, enabling faster and more efficient training.

Performance:

The model's performance improves as the number of MCMC iterations increases, and performance is enhanced by using data-based initialization for the Markov chains.

Scalability:

The approach is validated on synthetic tasks involving hypercubes and top-κ polytopes, proving its ability to recover true parameters in high-dimensional spaces.of the combinatorial size of the problem.) Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t-smooth in its first argument, and such that ∇θ lomegay (θ; y) = Ep(1) [Y] − y.

A similar result in the unconditional setting with data-based initialization is given in Proposition 6. In contrast, Sutskever and Tieleman (2010) showed that the expected CD-1 update with Gibbs sampling for restricted Boltzmann machines is not the gradient of any function, let alone a convex one.

3 Y 22 /t -

Improvements for AI systems

To improve AI systems using the methods proposed in Learning with Local Search MCMC Layers, I would implement the following specific architectural and algorithmic upgrades:

  1. Implement a Combinatorial MCMC Layer for end-to-end differentiable optimization. Instead of using standard continuous relaxations or expensive exact solvers, the system will use local search heuristics (like 2-opt, swap, or relocate) transformed into stochastic MCMC proposal distributions.

  2. Integrate Neighborhood Mixture Samplers to enable the model to learn from complex, multi-modal combinatorial spaces. This allows the AI to simultaneously explore different structural changes (e.g., swapping two items vs. relocating an entire sequence) within a single unified layer, preventing the model from getting stuck in local optima during training.

  3. Deploy Fenchel-Young Loss with Single-Step MCMC for high-speed training of structured prediction models. By utilizing the paper's mathematical proof that a specific target-dependent regularization allows for unbiased gradients even with just one MCMC iteration, I can significantly increase the training throughput of large-scale combinatorial models without sacrificing convergence quality.

  4. Apply Persistent/Data-Based MCMC Initialization to reduce gradient variance. By initializing the Markov chains at the ground-truth solution (for conditional tasks) or maintaining a persistent state (for unconditional tasks), the system will achieve faster convergence and higher precision in parameter estimation compared to standard Contrastive Divergence methods.


The improved AI system would be capable of:

  1. Solving large-scale, real-time dynamic logistics and routing problems (e.g., Dynamic Vehicle Routing with Time Windows) by predicting optimal prizes for service requests and refining them through a differentiable local search layer.

  2. Performing efficient structured prediction in complex discrete spaces where exact solutions are NP-hard, such as scheduling, network design, or molecular structure prediction.

  3. Training highly accurate generative models for combinatorial data (e.g., predicting binary vectors or subset selections) using significantly fewer computational resources and smaller datasets than traditional energy-based models.

  4. Scaling combinatorial optimization tasks to much larger problem instances by replacing the computational bottleneck of exact solvers with fast, differentiable, and stochastic local search layers.

Sources

Related papers