Salta al contenuto
PodcastTecnologiaMachine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)
Machine Learning Street Talk (MLST)
Ultimo episodio

264 episodi

  • Machine Learning Street Talk (MLST)

    Speech Recognition Is Not a Solved Problem — Pavan Muddireddy

    14/09/2026 | 1 h 42 min
    Pavankumar Reddy Muddireddy leads audio research at Mistral AI. He joins Tim Scarfe for a deep technical tour of Voxtral — and explains why the frontier of deployed voice is still a cascade of specialised models rather than one end-to-end system.

    IN PARTNERSHIP WITH MISTRAL AI:
    ---
    This episode was produced in partnership with Mistral AI.
    Mistral AI: https://mistral.ai/
    ---

    The conversation opens on architecture. Voxtral Chat feeds a 3B Ministral text trunk with continuous embeddings from an audio encoder, passed to the decoder as direct token input rather than through cross-attention as in Whisper, so the model can answer questions about emotion, timing and who spoke when without an intermediate transcript to lose them. The real-time model becomes a dual-stream decoder that consumes audio and emits text at once, at a target delay down to 160ms, with slower streams in parallel for anything that can wait for more context.

    On generation, Pavan explains why Voxtral TTS predicts continuous latents rather than discrete codec tokens, traces the lineage from SoundStream through EnCodec to Mimi's split of semantic and acoustic codebooks, and places FSQ and flow matching in it. Tim presses on the priors underneath: why a mel spectrogram instead of raw waveform, what noise augmentation buys, and when acoustic overfitting becomes somebody's fine-tuning problem. Then the failure modes. Diarisation is emitted autoregressively inside the transcript rather than by a separate head, which makes streaming diarisation fragile — less context, late speaker changes, invented extra speakers. And because the architecture commits to what it has already predicted, one out-of-distribution mistake compounds into looping or skipped segments, which is what DPO corrects: the negative supervision pre-training and SFT cannot give.

    The last third is the argument Tim keeps returning to. Customers running voice agents over millions of sessions describe scaffolding, not a solved problem, with a sharp drop outside the top few languages. Cascades survive because each component stays separately adaptable, observable and constrainable. And voice alone is cognitive debt: absorbing information and deciding in one serial stream is harder than glancing at a menu. Voice becomes ubiquitous beside a screen, not instead of one.

    ---
    TIMESTAMPS:
    00:00:00 Cold open
    00:00:46 Why Mistral moved into audio
    00:09:27 Inside Voxtral: trunk, encoder, dual streams
    00:20:22 Speech that works in real time
    00:30:52 How a voice becomes tokens
    00:39:59 Flow matching, FSQ and the new codec
    00:52:51 When speech models lose the speaker
    01:03:23 Correcting hallucinations with preferences
    01:12:12 Controlling synthetic speech
    01:20:06 Why cascades still win
    01:29:25 Speech in the wild
    01:33:46 Audio models as interfaces
    01:37:54 Why voice still needs a screen

    ---
    REFERENCES:
    paper:
    [00:01:42] Mistral 7B
    https://arxiv.org/abs/2310.06825
    [00:09:38] Voxtral
    https://arxiv.org/abs/2507.13264
    [00:14:41] Whisper: Robust Speech Recognition
    https://arxiv.org/abs/2212.04356
    [00:19:11] Voxtral Realtime
    https://arxiv.org/abs/2602.11298
    [00:21:52] Delayed Streams Modeling (Kyutai)
    https://arxiv.org/abs/2509.08753
    [00:30:52] Voxtral TTS
    https://arxiv.org/abs/2603.25551
    [00:32:38] SoundStream neural audio codec
    https://arxiv.org/abs/2107.03312
    [00:34:59] Flow Matching for Generative Modeling
    https://arxiv.org/abs/2210.02747
    [00:37:03] EnCodec: High Fidelity Neural Audio Compression
    https://arxiv.org/abs/2210.13438
    [00:37:42] Moshi and the Mimi codec
    https://arxiv.org/abs/2410.00037
    [00:39:05] Finite Scalar Quantization (FSQ)
    https://arxiv.org/abs/2309.15505
    [01:03:33] Direct Preference Optimization (DPO)
    https://arxiv.org/abs/2305.18290
    dataset:
    [00:46:14] Mozilla Common Voice
    https://commonvoice.mozilla.org/en/datasets
    organization:
    [00:50:47] Hugging Face
    https://huggingface.co/
  • Machine Learning Street Talk (MLST)

    How Replication Could Teach Machines What Good Science Looks Like — Edward Hughes

    11/09/2026 | 2 h 1 min
    Can a machine learn the judgement that separates a plausible-looking result from a faithful experiment? Edward Hughes, Chief Scientist and co-founder of Inherent, joins Tim Scarfe to argue that creativity is not optimisation, and that the missing capability in AI is choosing which questions are worth asking.

    SPONSOR:
    ---
    Cyber Fund built the Monastery to help founders ship products that were impossible a year ago.
    Apply now: https://cyber.fund
    ---

    Edward makes the case that Move 37 was innovative rather than creative, and that the field, not the individual, decides what counts as a discovery. That reframing runs through Csikszentmihalyi, Deutsch and exaptation into open-endedness, where deceptive goals and imperfect world models turn out to be the point rather than the problem. The second half turns to the paper: Replica, a task space built by redacting figures from real papers, and Faraday, a 27-billion-parameter model trained to steer a frontier coding agent that then beats the frontier on held-out replications.

    ---
    TIMESTAMPS:
    00:00:00 Cold open: Move 37, Faraday and collective intelligence
    00:01:08 Sponsor: CyberFund
    00:01:46 Inherent's $50M raise and the road from string theory
    00:09:14 Three timescales of learning: weights, context, culture
    00:13:47 Move 37 was innovative, not creative: the field decides
    00:20:39 Creativity as satisficing: the urinal and evolution
    00:25:06 Exaptation and the Tristan chord: creativity in context
    00:30:56 Coherence for whom? Deutsch's hard-to-vary explanations
    00:35:53 Why copying is creative: Deutsch and the constraint engineer
    00:42:27 Societies of agents and the strong Moravec paradox
    00:45:51 Evaluate in hindsight: from Lean proofs to climate change
    00:51:56 Picbreeder, local goals and why discovery needs deception
    00:57:21 Spaghetti proofs, translation layers and superhuman Go
    01:00:37 Does nature compress? Naturalness and real patterns
    01:07:36 Why replicate? Replica's redacted figures and Faraday
    01:12:31 Faraday beats Codex, Claude and GLM 5.2 on held-out tasks
    01:15:31 Replication to innovation: how the Transformer happened
    01:18:26 Deep replication: what Faraday learns from Voyager and GNoME
    01:23:37 Can the AI scientist cheat? Goodharting the judge
    01:29:09 Inside Replica: scale-down, 8xB300 runs, per-task rubrics
    01:34:11 The RL crisis: getting GRPO to work with per-turn credit
    01:39:43 Weights vs harnesses: AlphaEvolve, DGM and EvoTune
    01:45:45 The recursive company: agents cross a phase transition
    01:50:35 Collective intelligence and the electric dynamo
    01:55:46 What replaces OKRs? Incumbents and the burden of knowledge

    ---
    REFERENCES:
    MLST Creativity Article:
    https://archive.mlst.ai/read/why-creativity-cannot-be-interpolated

    organization:
    [00:01:47] Inherent
    https://inherentlabs.ai/
    other:
    [00:20:51] Marcel Duchamp, Fountain
    https://www.tate.org.uk/art/artworks/duchamp-fountain-t07573
    [00:05:19] Human-Timescale Adaptation in an Open-Ended Task Space (Adaptive Agent)
    https://arxiv.org/abs/2301.07608
    [00:06:05] The AI Scientist
    https://arxiv.org/abs/2408.06292
    [00:12:13] Training AI Scientists to Replicate Research (Replica and Faraday)
    https://arxiv.org/abs/2608.13331
    [01:44:46] Evolutionary Principles in Self-Referential Learning
    https://people.idsia.ch/~juergen/diploma.html
    [01:59:33] Are Ideas Getting Harder to Find?
    https://www.nber.org/papers/w23782
    book:
    [00:16:04] Creativity: Flow
    https://search.worldcat.org/title/254487436
    [00:26:22] Why Greatness Cannot Be Planned
    https://link.springer.com/book/10.1007/978-3-319-15524-1
    [00:33:03] The Beginning of Infinity
    https://www.penguinrandomhouse.com/books/293575/the-beginning-of-infinity-by-david-deutsch/
    [01:55:47] Laws of Knowledge
    https://www.penguin.co.nz/books/the-infinite-alphabet-9780241655672

    (Full list refs on YT/rescript)
    ---
    RESCRIPT:
    https://app.rescript.info/session/670296ba913761d0?share=6281911cac9bdbff637f10819d4d1e5c
  • Machine Learning Street Talk (MLST)

    AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen

    08/09/2026 | 1 h 29 min
    Could slowing AI development make superintelligence safer? Daniel Kokotajlo and Thomas Larsen of the AI Futures Project join Tim Scarfe to examine AI 2040: Plan A, a proposal to buy time before AI exceeds human control.

    SPONSOR:
    ---
    Cyber Fund built the Monastery to help founders ship products that were impossible a year ago.
    Apply now: https://cyber.fund
    ---

    After revisiting AI 2027 and the limits of forecasting, they ask what happens when AI can automate research and sustain an economy without human workers. Tim challenges the case for general models and asks whether intelligence alone explains power. Plan A proposes an initial pause to build safety infrastructure, then cautious development up to the strongest AI that can still be reliably controlled. The discussion tests the distinction between control and alignment, the case for public AI research, and whether the US and China could enforce a slowdown. It ends with the evidence that would change their forecasts.

    ---
    TIMESTAMPS:
    00:00:00 AI 2040: a slower route to superintelligence
    00:01:34 Sponsor: Cyber Fund
    00:02:12 From OpenAI to AI 2027
    00:06:58 Forecasts, war games and self-fulfilling prophecies
    00:17:44 Why AI sceptics are changing their minds
    00:23:04 When AI can replace its own researchers
    00:28:45 Could an AI economy grow without human workers?
    00:37:32 One general model or a society of specialists?
    00:47:43 Brains, machines and collective intelligence
    00:56:12 Plan A: buy time at the controllable frontier
    01:00:02 Why control buys time but cannot replace alignment
    01:06:36 Why AI research should be public
    01:10:32 Can the US and China enforce an AI slowdown?
    01:19:04 Why AI policy debates miss the technology
    01:21:56 Is AI normal technology? The remaining disagreement

    Many thanks to James Wilken-Smith for helping with show research.

    ---
    REFERENCES:
    other:
    [00:00:01] AI 2040: Plan A
    https://ai-2040.com/
    [00:03:27] AI 2027
    https://ai-2027.com/
    [00:13:47] Scenario Scrutiny for AI Policy
    https://blog.aifutures.org/p/scenario-scrutiny-for-ai-policy
    [00:33:11] The 2028 Global Intelligence Crisis
    https://www.citriniresearch.com/p/2028gic
    [01:00:40] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
    https://www.redwoodresearch.org/research/hugging-face-incident
    [01:09:21] The Hugging Face incident and the road ahead
    https://openai.com/index/hugging-face-incident-and-the-road-ahead/
    [01:22:01] AI as Normal Technology
    https://www.normaltech.ai/p/ai-as-normal-technology
    [01:22:51] Common Ground between AI 2027 & AI as Normal Technology
    https://asteriskmag.substack.com/p/common-ground-between-ai-2027-and
    person:
    [00:19:43] Geoffrey Hinton
    https://www.cs.toronto.edu/~hinton/
    [00:20:07] Ryan Greenblatt
    https://www.lesswrong.com/users/ryan_greenblatt
    [00:26:06] Elon Musk
    https://www.tesla.com/elon-musk
    tool:
    [00:21:46] ARC-AGI-3
    https://arcprize.org/arc-agi/3
    [00:21:53] AlphaGo and Move 37
    https://deepmind.google/research/alphago/
    [00:39:41] Claude
    https://claude.com/product/overview
    [00:39:58] NVIDIA H100 GPU
    https://www.nvidia.com/en-us/data-center/h100/
    paper:
    [00:24:42] Training AI Scientists to Replicate Research
    https://arxiv.org/abs/2608.13331v1
    [01:27:19] Validity of the single processor approach to achieving large scale computing capabilities
    https://www.cs.cmu.edu/~18742/papers/Amdahl1967.pdf
    book:
    [00:28:52] Bullshit Jobs: A Theory
    https://www.simonandschuster.com/books/Bullshit-Jobs/David-Graeber/9781501143335
    organization:
    [01:05:09] Redwood Research
    https://www.redwoodresearch.org/

    ---
    RESCRIPT:
    https://app.rescript.info/public/share/33d1a58fa8f307ae7dfd504d4fdaa9d5
  • Machine Learning Street Talk (MLST)

    Designing How AI Grows — Tom McGrath

    02/09/2026 | 1 h 40 min
    Tom McGrath is co-founder and Chief Scientist at Goodfire, and a former Google DeepMind researcher. He joins Tim Scarfe to ask what neural networks actually learn, whether their internal representations converge on structures in the world, and whether interpretability can extract new scientific knowledge rather than merely explain model outputs.

    Beginning with AlphaZero and learned modularity, the conversation moves into neural geometry: concept manifolds, reusable computation inside Llama, and why activation steering can fail when it pushes a model off-manifold. McGrath then makes the case for intentional design, using interpretability as part of the training loop. They examine controlled generalisation, features as rewards, predictive data debugging, and the uncomfortable fact that a model may recognise a hallucination or reward hack and still produce it.

    The discussion closes on grader awareness, oversight and collusion between adaptive agents, then returns to sparse autoencoders. SAEs are useful, McGrath argues, but they may fracture the higher-dimensional structures networks actually use. This episode was made with support from Goodfire.

    ---
    TIMESTAMPS:
    00:00:00 Introduction: Can interpretability speed-run science?
    00:02:03 The invisible grader
    00:06:51 What AlphaZero learned from the world
    00:12:24 Interpretability as a control loop
    00:21:54 The forbidden method and safer interventions
    00:37:36 Why models catch hallucinations too late
    00:46:19 Debug the dataset before training
    00:50:44 Why neural networks become modular
    00:55:57 Finding the geometry inside a network
    01:02:55 Why steering falls off the manifold
    01:12:10 A reusable calculator inside Llama
    01:17:19 From abstractions to goals
    01:25:28 Reward hacking, oversight and collusion
    01:37:23 Are sparse autoencoders dead?

    ---
    REFERENCES:
    paper:
    [00:05:45] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
    https://arxiv.org/abs/2502.17424v7
    [00:11:05] Acquisition of Chess Knowledge in AlphaZero
    https://arxiv.org/abs/2111.09259
    [00:25:30] Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
    https://arxiv.org/abs/2507.16795
    [00:29:30] Persona Vectors: Monitoring and Controlling Character Traits in Language Models
    https://arxiv.org/abs/2507.21509
    [00:41:14] Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
    https://arxiv.org/abs/2602.10067
    [00:47:03] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
    https://arxiv.org/abs/2606.12360
    [01:00:26] Do Sparse Autoencoders Capture Concept Manifolds?
    https://arxiv.org/abs/2604.28119
    [01:03:04] Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
    https://arxiv.org/abs/2605.05115
    [01:14:20] Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
    https://arxiv.org/abs/2605.01148
    [01:29:35] Measuring Reward-Seeking via Contrastive Belief Updates
    https://arxiv.org/abs/2607.18966v1
    other:
    [00:15:44] Intentional Design
    https://www.goodfire.com/blog/intentional-design
    [00:56:12] The World Inside Neural Networks
    https://www.goodfire.com/research/the-world-inside-neural-networks
    [01:37:28] A Pragmatic Vision for Interpretability
    https://www.alignmentforum.org/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-interpretability

    ---
    RESCRIPT:
    https://app.rescript.info/share/846cfee4131b664fd09209cc3b98018e
  • Machine Learning Street Talk (MLST)

    Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov

    22/08/2026 | 49 min
    Tim Scarfe speaks with Ilia Shumailov and Alexander Panfilov about their paper, Stealing Reasoning Traces from Proprietary LLM APIs.The core bug sounds deceptively simple: providers return encrypted reasoning state so conversations can be resumed or forked. But those blobs can be replayed across users and sibling models. A smaller model can ask the provider to decrypt the trace, then repeat the hidden reasoning in plain text. The discussion covers leaked private data, a broadly reusable jailbreak, poisoned agent traces, chain-of-thought monitoring, responsible disclosure, and possible defenses.Ilia Shumailov is an AI and security researcher, formerly at Google DeepMind, who completed his Cambridge PhD under Ross Anderson. Alexander Panfilov is a PhD researcher at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, working on AI safety, adversarial machine learning, and LLM red-teaming. They close by separating the demonstrated jailbreaking threat from ordinary benign distillation, and by arguing for controlled experiments over sweeping claims.---TIMESTAMPS:00:00:00 Intro montage00:01:33 Portable encrypted thought and decoded reasoning00:24:55 How the attack works and what it means00:39:04 Doom, defense, and scientific restraint---REFERENCES:paper:[00:00:00] Stealing Reasoning Traces from Proprietary LLM APIshttps://arxiv.org/abs/2608.09867[00:09:22] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyhttps://arxiv.org/abs/2507.11473[00:11:30] Reasoning Models Don’t Always Say What They Thinkhttps://www.anthropic.com/research/reasoning-models-dont-say-think[00:37:22] PostTrainBench: Can LLM Agents Automate LLM Post-Training?https://arxiv.org/abs/2603.08640[00:41:02] Large-scale online deanonymization with LLMshttps://arxiv.org/abs/2602.16800other:[00:09:28] OpenAI and Hugging Face partner to address security incident during model evaluationhttps://openai.com/index/hugging-face-model-evaluation-security-incident/[00:10:22] Claude, GPT, and Gemini All Struggle to Evade Monitorshttps://metr.org/notes/2025-08-22-claude-gpt-gemini-struggle-evade-monitors/tool:[00:42:08] Isabelle proof assistanthttps://isabelle.in.tum.de/---RESCRIPT: https://app.rescript.info/share/07fc38276e0823dc9b8986c32e202c7f
Altri podcast di Tecnologia
Su Machine Learning Street Talk (MLST)
Welcome! We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-depth analysis. Our approach is unrivalled in terms of scope and rigour – we believe in intellectual diversity in AI, and we touch on all of the main ideas in the field with the hype surgically removed. MLST is run by Tim Scarfe, Ph.D (https://www.linkedin.com/in/ecsquizor/) and features regular appearances from MIT Doctor of Philosophy Keith Duggar (https://www.linkedin.com/in/dr-keith-duggar/).
Sito web del podcast

Ascolta Machine Learning Street Talk (MLST), EasyApple e molti altri podcast da tutto il mondo con l’applicazione di radio.it

Scarica l'app gratuita radio.it

  • Salva le radio e i podcast favoriti
  • Streaming via Wi-Fi o Bluetooth
  • Supporta Carplay & Android Auto
  • Molte altre funzioni dell'app