Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

audscan

Find audio inside any binary file, extract it, and put edited sounds back.

Games usually keep their sounds as standard files packed inside their own archives: Wwise .wem, .bnk and .pck, FMOD sound banks, Ogg, plain WAV. audscan finds them by their headers, works out each one's exact size, and writes them out as ordinary files, without needing a tool for that particular engine.

It's a sibling of zscan (compressed streams) and texscan (textures), and works the same way: a scan writes a JSON manifest, and later steps work from it.

  • Scan a file for WAV, Wwise WEM (RIFF and big-endian RIFX), Wwise SoundBanks (.bnk) and file packages (.pck), FMOD FSB4 and FSB5 banks and Ogg streams, with their codec, channels, sample rate and length, and what's inside each bank or package.
  • Extract them as .wav, .wem, .bnk, .pck, .fsb and .ogg files, byte for byte. With --split, every sound inside a bank or package also comes out as a file of its own: Wwise WEMs named by their ID, FMOD tracks by their name.
  • Browse and play it all in a desktop app: a table of everything found, each bank's tracks, a waveform and playback.
  • Convert WEMs to files any player opens: Wwise Vorbis and Opus to Ogg (rewrapped, not re-encoded, so nothing is lost), PCM, IMA ADPCM and PTADPCM to WAV. audscan convert for WEM files, or extract --convert for everything extracted.
  • Put edited sounds back into a copy of the file: replace a whole file, a WEM inside a Wwise bank or package, or a track of an FMOD bank. Banks are rebuilt around the new sound and the result is checked before it's written.

audscan's desktop app: an FMOD bank's tracks, one selected with its waveform

Status: early. Scan, extract, splitting, WEM conversion, putting sounds back and the desktop app work. More formats are next.

Why audscan over vgmstream

vgmstream is the standard tool for playing game audio: it decodes hundreds of formats and plugs into foobar2000 and other players. It works on audio files you already have. audscan works a step earlier, and keeps the original data:

  • Finds audio anywhere. audscan scans any file for audio by its headers: packed archives, extracted blobs, memory dumps, formats nobody has written a reader for. It found 44,000 WEMs in Once Human's 81 GB of archives in a little over 2 minutes, without knowing its archive format.
  • Extracts the originals. Every file comes out byte for byte, as .wem, .bnk, .fsb or .ogg, with its exact size checked, so it can be modded, compared or put back.
  • Puts sounds back. Swap a WEM in a SoundBank or package, or a track in an FMOD bank, and audscan rebuilds the bank, fits it back into the game's file and checks the result. vgmstream only reads.
  • Splits banks into files you can use. Each sound in a Wwise bank or package becomes its own .wem, named by its ID and sorted by language. Each FMOD track becomes a WAV or a one-track .fsb that vgmstream and FMOD's tools still open.
  • Converts without re-encoding. Wwise Vorbis and Opus become standard Ogg files holding the same packets, a fraction of the size of the uncompressed WAV a decoder writes, with the audio unchanged.
  • Scriptable and safe. A JSON manifest, --json output, and batch extraction of whole games in one command. The input is never modified.
  • Tells you what's there. Every bank's contents with names, IDs, languages, codecs and lengths, plus notes on what's unusual, such as prefetch media that's only the start of a streamed sound.

When vgmstream is the better choice: to play or decode audio audscan can't. It handles far more codecs, including FMOD's Vorbis, console formats like XMA and ATRAC9, and loop points, and it plays inside your music player. The two work well together: audscan to find, extract and split, vgmstream to play anything audscan doesn't decode. audscan's own conversions are checked against vgmstream, and decode to exactly the same samples.

Download

Windows x64 builds are on the Releases page. The zip holds audscan.exe (command line) and audscan-gui.exe (desktop app). Nothing to install: the C runtime is built in.

Build

Rust 1.95 or later (1.89 for the command line alone):

cargo build --release

Usage

audscan scan game.pak -o manifest.json      # list audio, write a manifest
audscan scan game.pak --tracks              # also list what's inside each bank and package
audscan extract game.pak -m manifest.json -d audio/
audscan extract game.pak -d audio/ --split  # also each sound inside a bank or package
audscan extract game.pak -d audio/ --formats wem   # scan and extract in one go, WEMs only
audscan scan game.pak --formats fmod        # FSB4 and FSB5 only (and wwise: WEM, BNK, PCK)
audscan extract game.pak -d audio/ --split --convert   # every WEM also as .ogg or .wav
audscan convert audio/*.wem -d playable/    # convert WEM files you already have
audscan pack game.pak -m manifest.json -d audio/ -o game.new.pak   # put edited files back
      OFFSET        SIZE  FORMAT  CODEC               CH    RATE      LENGTH  NAME
        0xa4        4074  wav     PCM 16-bit           2   44100    0:00.023
      0x10bb        3094  wem BE  Wwise Vorbis         2   48000    0:10.000
      0x1cf0        1602  wem     Wwise Vorbis         1   32000    0:03.000
      0x3296        1450  fsb5    Vorbis               2   44100    0:12.500  3 tracks: music_intro, vo_line_01, amb_odd
      0x3cc7        1856  fsb4    mixed                2   44100    0:00.105  3 tracks: menu_theme, blip, voice_01
      0x4b48        1356  ogg     Vorbis               2   44100    0:02.268
      0x5da1        1994  bnk     mixed                1   48000    0:22.004  3 tracks: 111, 222, 333
      0x6812        1592  pck     Wwise Vorbis         2   48000    0:02.500  4 files (1 bnk, 3 wem): 777, 100, 100, ...
          #0         258  bnk     SoundBank            0       0           -  777 [sfx]
          #1         494  wem     Wwise Vorbis         2   48000    0:01.000  100 [sfx]
          #2         294  wem     Wwise Vorbis         1   48000    0:00.500  100 [english(us)]
          #3         344  wem     Wwise Vorbis         2   44100    0:01.000  4294967297 [sfx]
...

The rows starting # list what's inside a bank or package (--tracks).

Every command takes --json. The input is never modified. extract and pack refuse a file that no longer matches the manifest (--force overrides). --show-rejected lists headers that look like audio but can't be used, and why (a WAV with no fmt chunk, a file cut off by the end of the input...).

Desktop app

audscan-gui [FILE]

Open a file (or drop one on the window) and it's scanned straight away. Everything found is listed in a table you can sort and filter (by format, codec, name or offset), and a strip along the top shows where each file sits. Click one to see its details:

  • a bank or package lists its tracks (names or Wwise IDs, languages, codecs, lengths); pick one;
  • the waveform of the selected sound; Play (or Space) plays it, and clicking the waveform plays from there;
  • Save file… or Save track… writes it as it is; Save as Ogg/WAV… converts a WEM; Extract tracks… splits a whole bank.

Audio > Extract all… extracts everything, split and converted. Up and down arrows move through the list.

To put sounds back, pick Replace… next to a file or a track and choose the new one. It's checked straight away (the right kind of file, and whether it fits), marked as edited, and plays instead of the original (switch to the original to compare). Audio > Import edits from folder… takes every file you've changed in a folder written by Extract. The Pack window shows what's edited, does a dry run, and writes the packed file, which is read back and checked. The app asks before closing a file with edits that haven't been packed.

Packing a replaced WEM in the desktop app

It plays what audscan can decode: WAV, WEM (Vorbis, PCM, IMA ADPCM, PTADPCM) and Ogg Vorbis. Opus, FLAC and FMOD's own Vorbis aren't played yet (FMOD strips the Vorbis setup that decoding needs). Playback uses Windows' own audio output, so it's Windows-only; elsewhere the app works without it.

A WAV selected in the desktop app, with its waveform

The screenshots use made-up sounds from docs/make_demo.py.

Putting sounds back

audscan extract game.pak -m manifest.json -d audio/ --split        # 1. extract, with each bank's sounds
# 2. replace files in audio/ with your edited ones, keeping their names
audscan pack game.pak -m manifest.json -d audio/ --dry-run         # 3. see what would change
audscan pack game.pak -m manifest.json -d audio/ -o game.new.pak   # 4. write a new file

pack compares the folder with the original and puts back every file that changed. The input is never modified: the output is a new file (game.packed.pak by default), written to a temporary file, read back and checked before it's kept.

What was found Replace it with
A WAV, WEM, Ogg, FSB4 or FSB5 bank, SoundBank or package, whole a file of the same kind: a WEM for a WEM, a WAV for a WAV
A WEM or SoundBank inside a Wwise .bnk or .pck (split out as <ID>.wem, <ID>.bnk) a WEM (or SoundBank) in the same byte order
A track of an FSB5 bank (split out as <name>.fsb or .wav) a one-track .fsb of the bank's codec, from FMOD's FSBank or split out by audscan; for a PCM bank, also a WAV of its sample format

audscan doesn't encode audio: make a new WEM with Wwise (in the codec the game uses) and an FSB with FMOD's FSBank. It says so when a WEM's codec changes, since the game's sound objects may expect the old one.

How a changed sound fits:

  • Banks and packages are rebuilt. A SoundBank's media index and data are rewritten with the media moved to fit, keeping Wwise's 16-byte alignment, and a sound object that records the media's size (in HIRC) gets the new one. A package's lookup tables get the new sizes and positions, each file keeping its block alignment. An FSB5 bank gets the new track's header (with its own codec setup) and data, 32-byte aligned; the track keeps its name.
  • Smaller than the original: padded to keep its place, inside the file where its format allows (a RIFF JUNK chunk, a SoundBank's data section, an FSB5 bank's sample data), with zeros after it otherwise. Nothing else in the file moves.
  • At the end of the input (a .bnk, .pck or .wem on its own, or the FSB5 bank that ends an FMOD Studio .bank), it can grow or shrink: the output changes size, and size fields in front of it are updated: a RIFF header and chunk sizes that run to the end of the file, and an index entry giving its offset and size (the SNDH chunk of a .bank).
  • Bigger, in the middle of an archive, it doesn't fit, and pack says by how many bytes. The archive's own index would need rewriting, which audscan can't do for formats it doesn't know. Make the sound shorter, or pack the bank on its own if the game loads it as a separate file.

Not yet: prefetch media (replace the whole sound where it's streamed from), FSB4 tracks one by one (replace the whole .fsb), and growing a file in the middle of an archive.

Checked on real games, each result scanned again and played with vgmstream r2117:

  • Aniimo: in 12 SoundBanks, a WEM swapped for another from the same bank (Vorbis and Opus, bigger and smaller). Every WEM in each rebuilt bank is the intended one; vgmstream plays the new sound exactly as the replacement and the others as before.
  • BioShock Infinite: 3 WEMs replaced in the 486 MB SFX package, one growing from 388 KB to 10.9 MB. All 613 files in the new package are the intended ones, with their IDs and languages.
  • Slay the Spire 2: in a .bank cut out of the game, a track swapped for one 3.5 times its size. The RIFF, SNDH and SND sizes are updated, and vgmstream plays all 6 tracks, the new one exactly as its source.
  • Once Human: a WEM in the middle of a 287 MB archive replaced by a smaller one: padded with JUNK, nothing else in the archive changed, and it decodes exactly as the replacement.

Formats

WAV and Wwise WEM (RIFF, RIFX)

  • RIFF and big-endian RIFX files of form WAVE (and XWMA). Other RIFF forms (AVI, WebP, FMOD Studio .bank files) are skipped, so an FSB5 bank inside a .bank is still found.
  • The codec is named from the fmt chunk: PCM, IEEE float, A-law, mu-law, MS and IMA ADPCM, MP3, WMA, XMA/XMA2, ATRAC3/ATRAC9, WAVE_FORMAT_EXTENSIBLE, and Wwise's own (Vorbis, Opus, PTADPCM, IMA ADPCM, PCM).
  • Wwise audio is extracted as .wem: recognised by its codec, an akd chunk, being RIFX, or Wwise's own short fmt layout for ADPCM and PCM.
  • Lengths come from the data size (PCM, IMA ADPCM), the fact chunk, or the sample count Wwise stores for Vorbis, Opus and PTADPCM.
  • Common writer mistakes are tolerated, each with a note: a RIFF size that's 4 bytes short or counts a padding byte that isn't there, a last chunk a byte short, or a broken metadata chunk after the audio.

Checked against real games:

  • Once Human (Wwise): 81 GB of archives scanned in 140 seconds, finding 17,748 loose WEMs and 1,215 SoundBanks holding 26,500 more, every one sized exactly.
  • Aniimo and BioShock Infinite (Wwise, 2013 to now): all 637 loose WEMs, 2,022 SoundBanks and both packages scan to exactly their length. The 113,000 sounds inside (Vorbis, Opus, PTADPCM, IMA ADPCM, PCM) all get a codec and a length, and the lengths agree with each file's byte rate. All 10,200 WEMs split out of BioShock's packages are exact, voice lines in an english(us) folder.
  • Left 4 Dead 2: all 26,725 WAVs scan to exactly their length. 7,175 of them needed the writer-mistake handling above.
  • All 1,772 other .wav files on the development machine (Windows, KiCad, Unreal Engine samples) are exact too.
  • Against vgmstream r2117, on 1,944 complete WEMs (loose, and split from banks and packages): channels, rate and length agree for every Vorbis, Opus, IMA ADPCM and PCM file. PTADPCM lengths are the exact count Wwise stores; vgmstream counts whole blocks, which adds up to one block of padding (under 64 samples). A sample of 150 WEMs decodes to exactly its length.

Wwise SoundBanks (BNK) and file packages (PCK)

  • A SoundBank (BKHD...) is found whole, and its media index (DIDX) lists the WEMs in its DATA section: ID, codec, channels, rate and length of each. Banks without media (events only) are found too. Big-endian banks from older consoles work.
  • A file package (AKPK) is found whole, with every SoundBank, streamed WEM and external WEM (64-bit IDs) in its lookup tables, and each one's language from the package's own language map.
  • extract --split writes each of them as <ID>.wem or <ID>.bnk in a folder named after the bank or package, with localized files in a folder per language (the same ID is often used once per language). Scan a split-out .bnk to list what's inside it.
  • Banks often keep just the start of a streamed WEM ("prefetch" media) so it can start playing at once. Those are read from their header (codec, length) and noted as partial; the whole WEM is in the game's streamed files or .pck.

On Once Human, every one of the 1,215 banks was found without a problem. Of the WEMs split out of one archive, all 90 complete ones scan to exactly their length. 4 more are prefetch media and 4 banks hold media that aren't WEMs, likely plugin data such as reverb impulse responses (listed as unknown).

FMOD sound banks (FSB4, FSB5)

FSB5 is FMOD Studio's format (about 2013 on); FSB4 is FMOD Ex's (about 2006 to 2013).

  • The whole bank is found and extracted as one .fsb, with its track list: names, channels, sample rates, lengths and where each track's data is.
  • FSB4 banks with "basic headers" (only the first track described in full) work, and so does big-endian PCM. FSB4 doesn't record how the tracks' data is aligned; it's worked out from the data size.
  • extract --split writes each track as a file of its own, in a folder named after the bank:
    • PCM tracks (8-, 16-, 24-, 32-bit, float) as ordinary .wav files, converted where WAV needs it (signed 8-bit to unsigned, big-endian to little-endian).
    • Every other codec as a one-track .fsb of the same version: the bank's header cut down to that track, with its data copied unchanged. vgmstream, foobar2000 (with vgmstream) and FMOD's tools play these. The raw data alone wouldn't play: Vorbis in FSB5 has no setup headers, and XMA, ADPCM and the rest need their parameters from the track header.
    • Files are named after the track, with characters a file name can't hold replaced by _. Tracks with the same name get their number added, and unnamed ones are track<N>. Each file is read back and checked against the track before it's written.
  • All FSB5 codecs are named: PCM, GameCube ADPCM, IMA ADPCM, VAG/HEVAG, XMA, MPEG, CELT, ATRAC9, xWMA, Vorbis, FMOD ADPCM, Opus. FSB4 names each track's own (a bank can mix them): PCM, MPEG, IMA ADPCM, VAG, XMA, GameCube ADPCM, CELT.

FSB5 is checked against real banks. Slay the Spire 2 keeps FMOD Studio .bank files inside its Godot package: all 11 FSB5 banks are found (2,509 named Vorbis tracks, about 7 hours of audio), each ending exactly where its .bank does, and every split track is a one-track bank of exactly its length. So are the two in VTube Studio's Unity .resource file (Unity stores AudioClips as FSB5).

Split tracks play. Checked with vgmstream r2117: each of the 2,511 split tracks has the same length, rate, channels, codec and name as the same subsong of its original bank. 120 of them (226 MB of audio) decode to exactly the same samples either way. FSB4 is checked by test files only. Scanning 81 GB of a non-FMOD game's archives found no false FSB4 or FSB5 banks.

Ogg

  • Vorbis, Opus, FLAC and Speex, over any number of pages, each page's CRC checked.
  • Multiplexed streams (Theora video with Vorbis audio, say) are one file; chained files (one stream after another) are found as one file per stream.
  • A stream with no end-of-stream page is still found, with a note.

Converting WEMs

WEM codec Becomes How
Wwise Vorbis .ogg Vorbis headers rebuilt, packets rewrapped: no re-encoding
Wwise Opus .ogg Opus packets rewrapped in Ogg Opus: no re-encoding
PCM .wav copied
Wwise IMA ADPCM .wav decoded to 16-bit PCM
Wwise PTADPCM .wav decoded to 16-bit PCM, to the exact length Wwise stores
  • Wwise leaves out most of what a Vorbis decoder needs. audscan rebuilds the three Vorbis headers, unpacking the codebooks from Wwise's standard table (Wwise 2011.2 and later), and puts back the bits Wwise strips from each audio packet (as ww2ogg does). It also works out every page's granule position and trims the end to the exact sample count, so no separate revorb pass is needed.
  • Prefetch media (the first seconds of a streamed sound that a bank keeps) can't be converted on their own: the rest of the sound is in the game's streamed files or .pck. Split and convert the package to get the whole sound.
  • Vorbis with more than 2 channels keeps Wwise's channel order (as in WAV: L R C LFE...), which players read in Vorbis's order (L C R...). The channels can't be reordered without re-encoding, so the conversion notes it. Mono and stereo, nearly all game audio, are unaffected.

Checked with vgmstream r2117 on 1,690 real WEMs from Aniimo and BioShock Infinite: every converted file decodes to exactly the same samples as the WEM (520 Vorbis, 195 Opus, 275 IMA ADPCM including 40 stereo, PCM, and all 1,499 PTADPCM media in Aniimo's banks). The one exception is a 5.1 Vorbis track, whose samples are the same with the channels in Wwise's order, as above. The tests also check the rebuilt Vorbis headers with an independent decoder (lewton).

License

audscan is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 2 of the License, or (at your option) any later version (GPL-2.0-or-later). See LICENSE.

The Wwise Vorbis codebook table (crates/audscan-core/data/packed_codebooks_aoTuV_603.bin) comes from ww2ogg by Adam Gashlin, under the BSD 3-clause license in crates/audscan-core/data/COPYING-ww2ogg. The PTADPCM step table comes from vgmstream, under the ISC-style license in crates/audscan-core/data/COPYING-vgmstream. Both are compatible with the GPL.

About

Find, play, extract and convert audio inside game archives and other binary files: Wwise WEM, BNK and PCK, FMOD FSB4/FSB5 banks, WAV and Ogg. Splits banks into single sounds and converts WEMs to Ogg/WAV without re-encoding, with verification and a desktop app. A sibling of zscan and texscan.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages