Reinforcement learning for Quantum Tiq-Taq-Toe
cs.AI
Submitted: 2024-11-10
Updated: 2026-09-09
Code: https://github.com/Dinu23/Quantum-Tiq-Taq-Toe
License: http://creativecommons.org/licenses/by/4.0/
The gist: Quantum Tiq-Taq-Toe is a well-known benchmark and playground for both quantum computing and machine learning.
Terminology
Abstract
Quantum Tiq-Taq-Toe is a well-known benchmark and playground for both quantum computing and machine learning. Despite its popularity, no reinforcement learning (RL) methods have been applied to Quantum Tiq-Taq-Toe. Although there has been some research on Quantum Chess this game is significantly more complex in terms of computation and analysis. Therefore, we study the combination of quantum computing and reinforcement learning in Quantum Tiq-Taq-Toe, which may serve as an accessible testbed for the integration of both fields. Quantum games are challenging to represent classically due to their inherent partial observability and the potential for exponential state complexity. In Quantum Tiq-Taq-Toe, states are observed through Measurement (a 3x3 matrix of state probabilities) and Move History (a 9x9 matrix of entanglement relations), making strategy complex as each move can collapse the quantum state.
Sources
- Quantum Chess: Developing a Mathematical Framework and Design Methodology for Creating Quantum Games
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Proximal Policy Optimization Algorithms
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection