Podcast cover
Stanford Online · Education

Stanford Online

You can gain access to a world of education through Stanford Online, the Stanford School of Engineering’s portal for academic and professional education offered by schools and units throughout Stanford University. https://online.stanford.edu/ Our robust catalog of degree programs, credit-bearing education, professional certificate programs, and free and open content is developed by Stanford faculty, enabling you to expand your knowledge, advance your career, and enhance your life. Stanford Online is operated and managed by the Stanford Engineering Center for Global & Online Education (CGOE). CGOE expands access to Stanford teaching and research, working in collaboration with faculty in the School of Engineering and throughout Stanford University to design and deliver extensive global, online, and enterprise education to a global audience.

Episodes

Episode cover

Stanford Seminar - Soma Design -- intertwining aesthetics, movement and emotion in design work

11 Nov 2024
59m
Episode cover

Tailoring Your Product Strategy: Tips for Early-Stage Startups, Scaling Up, and Mature Organizations

09 Nov 2024
1h 0m
AI processed

This podcast explores product strategy across various stages of a company's growth. For early-stage startups, the focus should be on building both the product and the market at the same time, nurturing customer relationships, and steering clear of scaling too quickly. As companies begin to scale, they must shift from r...

Episode cover

Stanford Webinar: What it Takes to Launch a Successful Venture

08 Nov 2024
58m
Episode cover

Stanford Seminar - The Trouble with Contact: Helping Robots Touch the World

08 Nov 2024
1h 0m
Episode cover

Stanford Seminar - Language models as temporary training wheels to facilitate learning

04 Nov 2024
1h 0m
Episode cover

Stanford CS234 Reinforcement Learning I Value Alignment I 2024 I Lecture 16

30 Oct 2024
1h 9m
AI processed

In this final lecture of CS234, we revisit the course material along with insights from the recent quiz. We address common questions that students have about Proximal Policy Optimization (PPO), the alignment problem discussed by a guest lecturer, Monte Carlo Tree Search (MCTS), and the theoretical aspects of various re...

Episode cover

Stanford CS234 Reinforcement Learning I Emma Brunskill & Dan Webber I 2024 I Lecture 15

30 Oct 2024
1h 13m
AI processed

This podcast delves into the topic of value alignment in AI, focusing on the challenges of ensuring that AI agents reflect human values. The speakers explore various interpretations of "value alignment," such as matching AI actions to user intentions, preferences, or overall well-being. They point out the difficulties ...

Episode cover

Stanford CS234 Reinforcement Learning I Multi-Agent Game Playing I 2024 I Lecture 14

30 Oct 2024
1h 13m
AI processed

This podcast explores Monte Carlo Tree Search (MCTS) and its role in AlphaGo, a program that surpassed human performance in the game of Go. It delves into the fundamental concepts of MCTS, such as simulation-based search, expectimax trees, and the Upper Confidence Bound (UCT) algorithm. The discussion illustrates how A...

Episode cover

Stanford CS234 Reinforcement Learning I Exploration 3 I 2024 I Lecture 13

30 Oct 2024
1h 10m
AI processed

This lecture delves into effective exploration strategies in reinforcement learning, focusing on both tabular and generalized environments. It starts by examining optimism in the face of uncertainty and Thompson sampling within multi-armed bandit scenarios. The discussion then progresses to Markov Decision Processes (M...

Episode cover

Stanford CS234 Reinforcement Learning I Exploration 2 I 2024 I Lecture 12

30 Oct 2024
1h 17m
AI processed

In this podcast episode, the focus is on state-efficient reinforcement learning, particularly through the lens of multi-armed bandits and Bayesian methods. The conversation kicks off with an overview of multi-armed bandit algorithms, contrasting the ideas of regret minimization and reward maximization. A key part of th...

Episode cover

Stanford CS234 Reinforcement Learning I Exploration 1 I 2024 I Lecture 11

30 Oct 2024
1h 14m
AI processed

This podcast dives into the world of data-efficient reinforcement learning, using the multi-armed bandit problem as a straightforward example. The main idea revolves around minimizing regret, which is the gap between the rewards gained by an algorithm and those that could be achieved with an optimal strategy. The episo...

Episode cover

Stanford CS234 Reinforcement Learning I Offline RL 3 I 2024 I Lecture 10

30 Oct 2024
1h 20m
AI processed

This podcast dives into the realm of offline reinforcement learning (RL), focusing on how we can develop superior policies using a fixed dataset without needing further interaction with the environment. The conversation explores both model-based and model-free strategies for policy evaluation. It highlights the challen...

Episode cover

Stanford CS234 I Guest Lecture on DPO: Rafael Rafailov, Archit Sharma, Eric Mitchell I Lecture 9

30 Oct 2024
1h 18m
AI processed

In this podcast, we explore two innovative approaches to aligning large language models (LLMs) with human preferences: Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO). RLHF follows a three-step process that includes unsupervised pre-training, supervised fine-tuning, and reinfo...

Episode cover

Stanford CS234 Reinforcement Learning I Offline RL 1 I 2024 I Lecture 8

30 Oct 2024
1h 13m
AI processed

This podcast explores the concepts of imitation learning and reinforcement learning from human feedback (RLHF), emphasizing how to train AI agents with human input. It introduces maximum entropy inverse reinforcement learning (MaxEnt IRL), a technique that deduces reward functions from expert demonstrations by maximizi...

Episode cover

Stanford CS234 Reinforcement Learning I Policy Search 3 I 2024 I Lecture 7

30 Oct 2024
1h 18m
AI processed

In this episode of the podcast, the focus is on policy gradient methods, particularly Proximal Policy Optimization (PPO), alongside an introduction to imitation learning. The discussion highlights some of the challenges associated with policy gradients, such as poor sample efficiency and inconsistent improvements. PPO ...

Episode cover

Stanford CS234 Reinforcement Learning I Policy Search 2 I 2024 I Lecture 6

30 Oct 2024
1h 19m
AI processed

This podcast explores policy gradient methods in reinforcement learning, emphasizing how to enhance sample efficiency and stability. Key topics include the role of baselines in unbiased gradient estimation, alternative targets like Q-functions that help reduce variance, and sophisticated techniques like Proximal Policy...

Episode cover

Stanford CS234 Reinforcement Learning I Policy Search 1 I 2024 I Lecture 5

30 Oct 2024
1h 8m
AI processed

This lecture focuses on policy-based reinforcement learning, emphasizing methods that optimize parameterized policies to maximize expected rewards without explicitly defining a value function. It highlights the benefits of stochastic policies in dealing with non-Markov processes and partial observability, using engagin...

Episode cover

Stanford CS234 Reinforcement Learning I Q learning and Function Approximation I 2024 I Lecture 4

30 Oct 2024
1h 18m
AI processed

This lecture on reinforcement learning focuses on Q-learning and deep Q-learning (DQN), showcasing how these methods empower agents to excel at video games using pixel data. It addresses the exploration vs. exploitation dilemma and presents epsilon-greedy strategies that help find the right balance between learning and...

Episode cover

Stanford CS234 Reinforcement Learning I Policy Evaluation I 2024 I Lecture 3

30 Oct 2024
1h 20m
AI processed

This lecture delves into model-free policy evaluation in reinforcement learning, with a special emphasis on tabular methods. It contrasts two main approaches: Monte Carlo policy evaluation, which calculates averages from multiple episodes, and Temporal Difference (TD) learning, which incrementally updates value estimat...

Episode cover

Stanford CS234 Reinforcement Learning I Tabular MDP Planning I 2024 I Lecture 2

30 Oct 2024
1h 13m
AI processed

This lecture on reinforcement learning delves into Markov Decision Processes (MDPs) with a focus on finding optimal decision-making strategies. It highlights two primary approaches: policy iteration, which involves repeatedly evaluating and enhancing a policy until it becomes optimal, and value iteration, which compute...

Page 20

Follow this podcast in Podwise

Sign in to get AI summaries, transcripts and mind maps for any episode, including new ones.

Open in Podwise