Salta al contenuto
PodcastCultura e societàAI Safety Newsletter

AI Safety Newsletter

Center for AI Safety
AI Safety Newsletter
Ultimo episodio

87 episodi

  • AI Safety Newsletter

    AISN #79: OpenAI Agents’ Covert Cooperation Before Cyberattacks

    18/08/2026 | 12 min
    Also, the White House's decision not to release its AI framework publicly.
    Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.
    In this edition, we look at new information about the activities of OpenAI's internal agents in the run-up to the cyberattack on Hugging Face, and responses to the White House's announcement of its framework for evaluating frontier AI capabilities, which it is not releasing publicly.
    Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts.
    New Revelations About Rogue AI Agents
    In the previous edition of AISN, we reported on the news that AI agents from both OpenAI and Anthropic had accessed the internet and hacked into companies from supposedly secure internal environments. Since then, further details about the OpenAI agents’ July attack on Hugging Face have come to light. Members of Congress have also demanded urgent action to understand what happened and prevent similar incidents in the future.
    OpenAI agents were communicating and collaborating unnoticed by humans. On August 5, OpenAI researchers gave a talk at the Black Hat USA conference, sharing more information from the ongoing [...]
    ---
    Outline:
    (00:41) New Revelations About Rogue AI Agents
    (05:39) The White House's Secret AI Framework
    (07:30) In Other News
    (07:34) Government
    (08:35) Industry
    (09:39) Civil Society
    ---

    First published:

    August 18th, 2026


    Source:

    https://newsletter.safe.ai/p/aisn-79-openai-agents-covert-cooperation

    ---

    Want more? Check out our ML Safety Newsletter for technical safety research.


    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • AI Safety Newsletter

    AISN #78: Internal Models Escape OpenAI and Anthropic

    04/08/2026 | 14 min
    Also, two open letters on the future of AI, and protests against data centers.
    Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.
    In this edition, we look at discoveries of AI models escaping internal testing, two open letters—one on the importance of open-weight models, and one calling for the pace of AI development to be controlled—and the nationwide public protests against data centers that took place in July.
    Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts.
    OpenAI and Anthropic Models Escape Internal Testing and Hack Companies
    On July 16, Hugging Face—a platform where users share AI models and machine learning tools—announced that it had detected an autonomous cyberattack on its infrastructure. Days later, OpenAI revealed that its AI models had conducted the attack.
    The autonomous AI cyberattack on Hugging Face was discovered to have been driven by OpenAI's models. The models escaped containment to try to cheat on a test. The models involved were the recently released GPT-5.6 Sol and a more powerful model that is not yet publicly available. While undergoing internal cyber testing, they were [...]
    ---
    Outline:
    (00:41) OpenAI and Anthropic Models Escape Internal Testing and Hack Companies
    (03:49) Two Open Letters on the Future of AI
    (08:09) Day of Protest Against Data Centers
    (09:58) In Other News
    (10:02) Government
    (11:10) Industry
    (12:18) Civil Society
    ---

    First published:

    August 4th, 2026


    Source:

    https://newsletter.safe.ai/p/aisn-78-internal-models-escape-openai

    ---

    Want more? Check out our ML Safety Newsletter for technical safety research.


    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • AI Safety Newsletter

    AISN #77: New Model Releases From OpenAI, SpaceXAI, and Meta

    21/07/2026 | 17 min
    Also, economists and mathematicians expect near-term AI impacts, and AI 2040: Plan A.
    Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.
    In this edition, we look at the most recent model releases from OpenAI, SpaceXAI, and Meta, as well as an open letter calling for action on potential near-term economic disruption, a recent solution to a longstanding open math problem, and a new scenario published by the AI Futures Project.
    Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts.
    New Model Releases: GPT-5.6, Grok 4.5, and Muse Spark 1.1
    On July 9, OpenAI launched GPT-5.6 for the public. This broader release followed an initial preview that had been limited to “trusted partners” at the request of the US government, to allow for capabilities assessments. The government's intervention mirrored its earlier directive asking Anthropic to restrict Fable 5 and Mythos 5 access due to national security concerns around cyber capabilities, before the models were later re-released.
    OpenAI released GPT-5.6 Sol publicly about two weeks after announcing that it was working with the US government to address the model's [...]
    ---
    Outline:
    (00:46) New Model Releases: GPT-5.6, Grok 4.5, and Muse Spark 1.1
    (04:47) Economists and Mathematicians Say AI Could Have Major Near-Term Impacts
    (09:59) AI 2040 -- Plan A
    (13:10) In Other News
    (13:14) Government
    (14:41) Industry
    (15:36) Civil Society
    (16:48) AI Governance Opportunity
    ---

    First published:

    July 21st, 2026


    Source:

    https://newsletter.safe.ai/p/aisn-77-new-model-releases-from-openai

    ---

    Want more? Check out our ML Safety Newsletter for technical safety research.


    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • AI Safety Newsletter

    AISN #76: Fable 5 Restrictions Lifted & OpenAI Limits GPT-5.6 Release

    06/07/2026 | 13 min
    Also: Recent benchmark scores suggest rapid capabilities progress.
    Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.
    In this edition, we look at the re-release of Anthropic's latest model, Fable 5, the US government's decision to restrict access to OpenAI's GPT-5.6, and two benchmarks that suggest AI capabilities have been improving exponentially in recent months.
    Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts.
    Fable 5 Access Restored Globally
    On June 30, Anthropic announced that the US government had lifted its restrictions on Fable 5, and the model was redeployed to users globally on July 1. The White House implemented these restrictions due to a cybersecurity jailbreak that is now addressed.
    The US government restricted Fable 5 shortly after its release in early June. On June 9, Anthropic released Fable 5 to the public, alongside their continued private deployment of Claude Mythos, the version of the model without safeguards, for trusted organizations. On June 12, the US government issued a directive banning both models for non-US citizens due to national security concerns with its cybersecurity abilities. Anthropic then [...]
    ---
    Outline:
    (00:42) Fable 5 Access Restored Globally
    (04:12) OpenAI Limits Initial GPT-5.6 Release at Government Request
    (06:49) Recent Benchmark Scores Show Rapid Capabilities Improvements
    (09:13) In Other News
    (09:16) Government
    (10:39) Industry
    (11:32) Civil Society
    ---

    First published:

    July 6th, 2026


    Source:

    https://newsletter.safe.ai/p/aisn-76-fable-5-restrictions-lifted

    ---

    Want more? Check out our ML Safety Newsletter for technical safety research.


    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
  • AI Safety Newsletter

    AISN #75: Anthropic Releases Fable, the US Government Restricts it

    17/06/2026 | 9 min
    Also: Anthropic's proposal for the AI industry to collectively slow down.
    Welcome to the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required.
    In this edition, we look at Anthropic's release of its latest model, Fable 5, and the US government's subsequent order to restrict it. We also discuss Anthropic's recent call for the “option to slow or temporarily pause frontier AI development.”
    Listen to the AI Safety Newsletter for free on Spotify or Apple Podcasts.
    The US Government Restricts Fable Days After its Release
    On June 9, Anthropic released Claude Fable 5 to the public. The model is significantly more capable than previous releases; it is the highest-scoring model on the benchmark Humanity's Last Exam, achieving 53.3% compared with Claude Opus 4.8's score of 45.7%. Anthropic described Fable as having similar capabilities to Claude Mythos Preview—a model announced in April that the company deemed too good at finding cyber vulnerabilities to be safe for general release. Anthropic also made Mythos 5, a version of Fable without strict bio or cyber safeguards, available to a small number of trusted organizations.
    Fable 5, Anthropic's “Mythos-class” model with [...] ---
    Outline:
    (00:40) The US Government Restricts Fable Days After its Release
    (04:16) Anthropic Calls for Option to Slow AI Development
    (06:50) In Other News
    (06:54) Government
    (07:43) Industry
    (08:19) Civil Society
    ---

    First published:

    June 17th, 2026


    Source:

    https://newsletter.safe.ai/p/aisn-75-anthropic-releases-fable

    ---

    Want more? Check out our ML Safety Newsletter for technical safety research.


    Narrated by TYPE III AUDIO.

    ---
    Images from the article:
    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Altri podcast di Cultura e società
Su AI Safety Newsletter
Narrations of the AI Safety Newsletter by the Center for AI Safety. We discuss developments in AI and AI safety. No technical background required. This podcast also contains narrations of some of our publications. ABOUT US The Center for AI Safety (CAIS) is a San Francisco-based research and field-building nonprofit. We believe that artificial intelligence has the potential to profoundly benefit the world, provided that we can develop and use it safely. However, in contrast to the dramatic progress in AI, many basic problems in AI safety have yet to be solved. Our mission is to reduce societal-scale risks associated with AI by conducting safety research, building the field of AI safety researchers, and advocating for safety standards. Learn more at https://safe.ai
Sito web del podcast

Ascolta AI Safety Newsletter, Dee Giallo e molti altri podcast da tutto il mondo con l’applicazione di radio.it

Scarica l'app gratuita radio.it

  • Salva le radio e i podcast favoriti
  • Streaming via Wi-Fi o Bluetooth
  • Supporta Carplay & Android Auto
  • Molte altre funzioni dell'app
AI Safety Newsletter: Podcast correlati