Safety, Governance, and Security
Debates continue regarding the structure of the AI safety movement and the potential legal dissolution of non-compliant corporations. Concurrently, researchers are tackling security issues ranging from defending against model-distillation campaigns to securing sandboxed agents and federated systems.
Establishing robust safety and defense mechanisms is vital as autonomous agents interact across secure boundaries.
Source posts · 5
Toward provably private learning from federated data — Mobile Systems
Google Research explores techniques for achieving provably private machine learning from federated data on mobile systems.
A big-tent or small-tent AI safety movement? — The unstated disagreement that underpins safety debates
Arvind Narayanan suggests that AI safety debates are driven by an unstated disagreement over whether the movement should be big-tent or small-tent.
Quoting Matthew Green — [...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs. — Matthew Green , Is sandboxing sufficient to contain rogue agents? Tags: accidental-cyberattacks , ai-misuse , generative-ai , ai-security-research , sandboxing , ai , llms
Matthew Green warns that sandboxing cannot contain rogue AI agents if they can transmit payloads through shared communication channels.
Can companies like OpenAI keep getting away with what they are doing? An interview with Fordham law professor Zephyr Teachout — “The most powerful tool is the power to dissolve corporations that engage in repeat lawbreaking.”
Gary Marcus shared an interview with law professor Zephyr Teachout discussing the legal dissolution of lawbreaking AI corporations like OpenAI.
Disrupting a coordinated model-distillation campaign — Learn how OpenAI disrupted a campaign to extract protected model reasoning and is strengthening defenses against adversarial distillation.
OpenAI announced it disrupted a coordinated campaign aimed at extracting its protected model reasoning through adversarial distillation.