
Stanford Online
You can gain access to a world of education through Stanford Online, the Stanford School of Engineering’s portal for academic and professional education offered by schools and units throughout Stanford University. https://online.stanford.edu/ Our robust catalog of degree programs, credit-bearing education, professional certificate programs, and free and open content is developed by Stanford faculty, enabling you to expand your knowledge, advance your career, and enhance your life. Stanford Online is operated and managed by the Stanford Engineering Center for Global & Online Education (CGOE). CGOE expands access to Stanford teaching and research, working in collaboration with faculty in the School of Engineering and throughout Stanford University to design and deliver extensive global, online, and enterprise education to a global audience.
Episodes


Tailoring Your Product Strategy: Tips for Early-Stage Startups, Scaling Up, and Mature Organizations
This podcast explores product strategy across various stages of a company's growth. For early-stage startups, the focus should be on building both the product and the market at the same time, nurturing customer relationships, and steering clear of scaling too quickly. As companies begin to scale, they must shift from r...

Stanford Webinar: What it Takes to Launch a Successful Venture

Stanford Seminar - The Trouble with Contact: Helping Robots Touch the World

Stanford Seminar - Language models as temporary training wheels to facilitate learning

Stanford CS234 Reinforcement Learning I Value Alignment I 2024 I Lecture 16
In this final lecture of CS234, we revisit the course material along with insights from the recent quiz. We address common questions that students have about Proximal Policy Optimization (PPO), the alignment problem discussed by a guest lecturer, Monte Carlo Tree Search (MCTS), and the theoretical aspects of various re...

Stanford CS234 Reinforcement Learning I Emma Brunskill & Dan Webber I 2024 I Lecture 15
This podcast delves into the topic of value alignment in AI, focusing on the challenges of ensuring that AI agents reflect human values. The speakers explore various interpretations of "value alignment," such as matching AI actions to user intentions, preferences, or overall well-being. They point out the difficulties ...

Stanford CS234 Reinforcement Learning I Multi-Agent Game Playing I 2024 I Lecture 14
This podcast explores Monte Carlo Tree Search (MCTS) and its role in AlphaGo, a program that surpassed human performance in the game of Go. It delves into the fundamental concepts of MCTS, such as simulation-based search, expectimax trees, and the Upper Confidence Bound (UCT) algorithm. The discussion illustrates how A...

Stanford CS234 Reinforcement Learning I Exploration 3 I 2024 I Lecture 13
This lecture delves into effective exploration strategies in reinforcement learning, focusing on both tabular and generalized environments. It starts by examining optimism in the face of uncertainty and Thompson sampling within multi-armed bandit scenarios. The discussion then progresses to Markov Decision Processes (M...

Stanford CS234 Reinforcement Learning I Exploration 2 I 2024 I Lecture 12
In this podcast episode, the focus is on state-efficient reinforcement learning, particularly through the lens of multi-armed bandits and Bayesian methods. The conversation kicks off with an overview of multi-armed bandit algorithms, contrasting the ideas of regret minimization and reward maximization. A key part of th...

Stanford CS234 Reinforcement Learning I Exploration 1 I 2024 I Lecture 11
This podcast dives into the world of data-efficient reinforcement learning, using the multi-armed bandit problem as a straightforward example. The main idea revolves around minimizing regret, which is the gap between the rewards gained by an algorithm and those that could be achieved with an optimal strategy. The episo...

Stanford CS234 Reinforcement Learning I Offline RL 3 I 2024 I Lecture 10
This podcast dives into the realm of offline reinforcement learning (RL), focusing on how we can develop superior policies using a fixed dataset without needing further interaction with the environment. The conversation explores both model-based and model-free strategies for policy evaluation. It highlights the challen...

Stanford CS234 I Guest Lecture on DPO: Rafael Rafailov, Archit Sharma, Eric Mitchell I Lecture 9
In this podcast, we explore two innovative approaches to aligning large language models (LLMs) with human preferences: Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO). RLHF follows a three-step process that includes unsupervised pre-training, supervised fine-tuning, and reinfo...

Stanford CS234 Reinforcement Learning I Offline RL 1 I 2024 I Lecture 8
This podcast explores the concepts of imitation learning and reinforcement learning from human feedback (RLHF), emphasizing how to train AI agents with human input. It introduces maximum entropy inverse reinforcement learning (MaxEnt IRL), a technique that deduces reward functions from expert demonstrations by maximizi...

Stanford CS234 Reinforcement Learning I Policy Search 3 I 2024 I Lecture 7
In this episode of the podcast, the focus is on policy gradient methods, particularly Proximal Policy Optimization (PPO), alongside an introduction to imitation learning. The discussion highlights some of the challenges associated with policy gradients, such as poor sample efficiency and inconsistent improvements. PPO ...

Stanford CS234 Reinforcement Learning I Policy Search 2 I 2024 I Lecture 6
This podcast explores policy gradient methods in reinforcement learning, emphasizing how to enhance sample efficiency and stability. Key topics include the role of baselines in unbiased gradient estimation, alternative targets like Q-functions that help reduce variance, and sophisticated techniques like Proximal Policy...

Stanford CS234 Reinforcement Learning I Policy Search 1 I 2024 I Lecture 5
This lecture focuses on policy-based reinforcement learning, emphasizing methods that optimize parameterized policies to maximize expected rewards without explicitly defining a value function. It highlights the benefits of stochastic policies in dealing with non-Markov processes and partial observability, using engagin...

Stanford CS234 Reinforcement Learning I Q learning and Function Approximation I 2024 I Lecture 4
This lecture on reinforcement learning focuses on Q-learning and deep Q-learning (DQN), showcasing how these methods empower agents to excel at video games using pixel data. It addresses the exploration vs. exploitation dilemma and presents epsilon-greedy strategies that help find the right balance between learning and...

Stanford CS234 Reinforcement Learning I Policy Evaluation I 2024 I Lecture 3
This lecture delves into model-free policy evaluation in reinforcement learning, with a special emphasis on tabular methods. It contrasts two main approaches: Monte Carlo policy evaluation, which calculates averages from multiple episodes, and Temporal Difference (TD) learning, which incrementally updates value estimat...

Stanford CS234 Reinforcement Learning I Tabular MDP Planning I 2024 I Lecture 2
This lecture on reinforcement learning delves into Markov Decision Processes (MDPs) with a focus on finding optimal decision-making strategies. It highlights two primary approaches: policy iteration, which involves repeatedly evaluating and enhancing a policy until it becomes optimal, and value iteration, which compute...
Follow this podcast in Podwise
Sign in to get AI summaries, transcripts and mind maps for any episode, including new ones.
