Q-Learning Lab is a lightweight, dependency-free educational tool designed to demystify the tabular Q-learning algorithm for undergraduate students. While traditional gridworld visualizers often hide the underlying mechanics of reinforcement learning behind animations, this tool prioritizes transparency. It provides a live, step-by-step breakdown of the Bellman update equation and allows learners to export a complete, granular trace of their agent's decision-making process as a CSV file for independent analysis.
The Learn-Export-Analyze Loop
The tool is built around a constructionist pedagogical framework. Instead of merely watching an agent converge, students engage in a three-part cycle:
Learn: Students run an agent in a 5x5 gridworld, observing the Bellman substitution panel to see exactly how Q-values are updated in real-time.
Export: The tool logs every transition, including the pre-action Q-row, the greedy-versus-random decision, and wall-collision events, which can be exported as a CSV.
Analyze: Students use the exported data to create their own visualizations, such as learning curves, value heatmaps, and state-visitation maps, turning the agent's history into a dataset for reflective inquiry.
Why It Matters
By making the algorithm's internal state fully inspectable, Q-Learning Lab addresses a common pedagogical gap: the disconnect between the abstract Bellman equation and the agent's observed behavior. The tool's design—a single, bilingual (Thai/English) HTML file—removes technical barriers to entry, making it accessible for diverse classroom environments. By treating the agent's trace as an artifact for analysis, the tool encourages students to think like researchers, using evidence to diagnose exploration failures or reward misspecifications rather than relying on intuition alone.
A real-time interface component that displays the specific numeric values used in the Q-learning update equation at each step.
Decision-complete trace
A comprehensive log of every agent transition, including the full Q-row, the decision-making logic (greedy vs. random), and the resulting temporal-difference error.
Learn-export-analyze loop
A pedagogical cycle where students generate data through agent training, export it, and perform their own quantitative analysis to understand learning dynamics.
Tabular Q-learning
A reinforcement learning algorithm that maintains a table of values for every state-action pair to estimate the expected future reward.
Off-policy
A property of Q-learning where the agent can learn the optimal policy even while following a different behavior policy, such as one driven by exploration.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.