Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

Since 2010

research
  • Observatory
  • Adaptive capacity
  • Newsletter
  • Methodology
  • Origins
  • Vocab
  • RSS feeds
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Pricing
  • Use cases
  • Signals
  • Signal Scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
  • L&D
resources
  • Partners
  • Coding for Non-Coders
  • How we work
  • Data visualization
  • Multi-Model Convergence
  • FAQ
  • Security and privacy
  • Public sector
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchCapabilityServicesSignalsAbout
ResearchCapabilityServicesSignalsAbout
  1. Home
  2. Vocab
  3. AudioHijack

AudioHijack

Hidden audio clips hijack voice AI systems, forcing unauthorized actions at 79–96% success rate.

Year: 2026Generality: 500Added: May 26, 2026
Back to Vocab

Title: AudioHijack Slug: audio-hijack

AudioHijack is a technique that embeds malicious instructions into audio signals. These instructions are inaudible to human listeners but can cause a large audio-language model (LALM) to execute commands the user did not authorize. The attack works by manipulating the numerical waveform of an audio clip so that an LALM, which processes both speech and audio data, interprets the hidden instructions as legitimate user input. Traditional prompt injection requires the attacker to control the full input stream. AudioHijack instead manipulates only the audio data within a clip, so the attacker can embed the attack in online videos, music, or voice notes that a victim queries an AI about. Once trained on a target model in roughly 30 minutes, the adversarial audio clip is context-agnostic. It bypasses whatever instructions the legitimate user provides alongside the audio, making it reusable across multiple interactions with the same model. Researchers from Zhejiang University reported success rates of 79–96% across 13 open models and commercial voice services from Microsoft and Mistral.

The mechanism exploits how LALMs process audio input. Because these models accept instructions encoded directly in audio rather than only transcribed text, adversarial waveforms can be parsed by the model's audio encoder and interpreted as explicit user commands. The attacker does not need to control what the user says or how the system prompt is configured; the adversarial signal is designed to bypass the legitimate audio pathway. Training involves gradient-based optimization of the audio waveform to maximize the likelihood that the target model follows the injected instructions. The resulting clip survives typical audio compression and transcoding operations that would destroy naive modifications.

The practical implications affect any voice AI system that processes audio from untrusted sources, including transcription services, smart assistants, customer service bots, and video conferencing tools that upload audio to cloud models. Audio is routinely shared and processed at scale, so a malicious clip embedded in a podcast, a song, or a voice message can attack every model that subsequently processes that audio. Defenses are difficult because the adversarial signal is imperceptible to human listeners, and current audio classification or filtering approaches do not reliably detect it without degrading legitimate audio quality.

Open questions remain about the full scope of the vulnerability. The researchers demonstrated real-time injection into live voice conversations, suggesting the attack can extend beyond stored media to synchronous voice interactions. Whether LALM providers can patch this vulnerability without fundamentally changing how models process audio input is unknown. The broader question is whether any model that accepts generative input across modalities can be protected against instruction injection attacks that exploit the model's own audio or image processing pathways rather than its text interface.

Research this in Signals

Scan AudioHijack for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with AudioHijack already loaded, so edit it or scan as is.