Everything M2Suite knows about 3DO and Panasonic M2 data, written down so the knowledge survives independently of the code. Each entry says what the layout is, how it was established, and — where it matters — which published description is wrong.
Conventions: u8/u16/u32 unsigned, s16 signed, [n] array. 3DO M1 and M2
chunk containers are big-endian; Alone in the Dark data is
little-endian (it is a DOS port). Byte offsets are from the start of the
structure being described.
Contents
- Reading this document
- IFF-style containers
- Textures — UTF / M2TX
- Cels — CCB / PDAT / PLUT
- ANIM and cel chains
- IMAG screens and LRForm
- RSRC resource bundles
- Audio codecs
- DataStreamer movies
- Standalone FILM
- Cinepak
- MPEG-1 in 3DO containers
- Opera filesystem
- M2 PowerPC executables
- Alone in the Dark
- Formats that resisted
Findings are graded, because not all of them are equally certain:
- Verified — decoded output was checked against something independent (a reference implementation, a known-good render, a statistical test, or an exact file-size match).
- Consistent — the layout explains every file we have, but nothing external corroborates it.
- Assumed — a working guess that has not failed yet.
3DO and M2 both use FORM-style chunk containers, but with different
padding rules, and the difference is not cosmetic.
u32 tag four-character chunk id
u32 size chunk size INCLUDING these 8 header bytes
u8 payload[size - 8]
u8 pad[] to the alignment boundary
Finding (verified): M2 chunks align to 4 bytes; classic IFF-85
chunks align to 2. A parser that assumes one alignment silently
desynchronises on the other after the first odd-sized chunk. libm2core's
IffForm accepts both and picks per container.
Container ids seen in the wild: FORM, CAT , LIST. A CAT behaves as
a concatenation of FORMs and must be walked, not treated as a leaf.
M2's native texture format. FORM+TXTR, or CAT +TXTR for a bundle of
several textures in one file.
Sub-chunks:
| Chunk | Contents |
|---|---|
M2TX |
Header: dimensions, texel format, LOD count, flags |
M2PI |
PIP — the colour lookup table (palette) |
M2CI |
DCI — colour/alpha component description |
M2TD |
The texel data itself, LOD by LOD |
M2LR |
LOD run lengths for compressed texel data |
Finding (consistent): LOD 0 is the largest mip level and appears first. Texel data for later LODs follows contiguously; the per-LOD sizes come from the header's dimensions, halved and floored per level, not from a stored table.
Currently decoded: PIP-indexed, uncompressed, 8-bits-per-index. Other texel formats (RLE-compressed runs, direct/non-palette formats) are detected and rejected rather than guessed at — see LIMITATIONS.md.
Ground truth for this format is the SDK Mercury texture library C source; see REFERENCES.md.
Finding (verified). Some 3DO Gamepack art wraps a normal cel in an
OFST offset-table header rather than starting with CCB directly:
'OFST' u32 celOffset u32 sibling[...]
... the real 'CCB ' chunk begins at celOffset
celOffset points straight at the CCB; the remaining u32s index sibling
resources. Skipping to the CCB decodes the image normally — verified against
starC.cel (40×41), ring3.cel (320×240), shadow.cel (16×16) and
AJLbanner.cel (a 320×240 title banner). The loader unwraps OFST
transparently.
The 3DO M1 sprite/bitmap primitive. A cel is a chunk chain:
| Chunk | Contents |
|---|---|
CCB |
Cel control block: flags, dimensions, bit depth, pixel-data offsets |
PDAT |
The pixel data |
PLUT |
Palette lookup table (16-bit 0RRRRRGGGGGBBBBB entries) |
XTRA |
Optional extra description |
Bit depths 1, 2, 4, 6, 8 and 16 all occur. Depths below 8 are coded (palette-indexed via the PLUT); 16-bit is direct colour.
Finding (verified). An 8bpp coded pixel is not simply a palette index. The low 5 bits index the 32-entry PLUT; the top 3 bits are the AMV, a per-pixel brightness multiplier that the cel engine's pixel processor applies as it draws. Masking them off — the obvious reading — throws away all of a sprite's shading and renders it posterised.
The CCB's PIXC word holds two 16-bit PPMPC values: PPMP_0 in the low
half for pixels whose SSB is clear, PPMP_1 in the high half for pixels
with it set. For a coded cel the SSB is bit 15 of the PLUT entry. Field
layout, from 3DO Portfolio 2.5 Interfaces/2p5/includes/hardware.h:
| Bits | Field | Meaning |
|---|---|---|
| 15 | 1S |
first source: 0 = cel pixel, 1 = framebuffer |
| 14–13 | MS |
multiplier source: 0 = the CCB's MF, 1 = the pixel's AMV (PPMPC_MS_PIN) |
| 12–10 | MF |
multiply factor, stored as value−1 (0b111 = ×8) |
| 9–8 | SF |
divide factor — 0b00 = 16, 01 = 2, 10 = 4, 11 = 8 |
| 7–6 | 2S |
second source: 0 = none, 1 = the CCB's AV, 2 = framebuffer |
| 5–1 | AV |
add value |
| 0 | 2D |
second-source doubling |
Two traps worth naming. PPMPC_SF_16 is 0x0000, so a zero SF field
means divide by sixteen, not by one. And the AMV is encoded the same way
MF is — as value−1 — so a fully lit pixel (AMV 7) multiplies by 8 and,
with the usual divide by 8, reproduces the palette colour exactly. Reading
the AMV as a raw 0–7 caps every pixel at 7/8 brightness and crushes AMV-0
pixels, roughly a third of a typical sprite, to pure black.
Verified against Escape from Monster Manor's zombie.cels: PIXC is
0x3f803f01, so PPMP_0 (all 32 PLUT entries have bit 15 clear) is
MS_PIN with SF = 8 and no second source — a pure colour × (AMV+1) / 8.
Its AMV values span the full 0–7 range across 4552 opaque pixels. The
bundled Message.cel fixture uses the same mode (0x3f003f00).
M2Suite does not model the second-source term: it reads the framebuffer, which a still decode does not have. Sprites that use it are doing blend effects, not carrying colour.
The streaming Cinepak output cel needs no such correction. The 3DO's
CPakSubscriber.c builds its LRForm CCB with ccb_PIXC = 0x1F811F00, and
both halves work out to ×8/8 — unity. Only the second source differs
(PPMP_1 adds the framebuffer). Decoded video frames are passed straight
through at their own brightness.
Packed (RLE) cels store rows as a chain: each row begins with a length word, then a run of control bytes selecting literal / repeat / transparent spans.
Finding (verified): 1-bit packed cels use a variable-length row
preamble — the offset from the row start to the first control byte is not
fixed. Guessing it produces plausible-looking garbage. The reliable test is
to validate the row chain: walk all rows using a candidate preamble length
and accept it only if every row lands exactly on the next row's start and
the last row ends at the end of the data. Verified against Street Fighter's
title screens and StarBlade's DANGER label.
Some files are raw pixel data with no CCB . These are M1 VRAM captures
and are stored in LRForm layout: the framebuffer is split into two
interleaved halves (even and odd scanlines in separate regions), so the
image must be de-interleaved before it looks like anything.
Finding (verified). form3do.h describes AnimChunk::frameRate as the
"number of 1/60s of a sec to display each frame" and types it int32, which
reads like a plain integer. It is not — it is 16.16 fixed point, so the
count of 60ths is frameRate / 65536 and the rate is
60 / (frameRate / 65536) fps.
Confirmed across 350+ ANIM files on real discs: every value is a clean 16.16 number between 2.0 and 20.0, and the fractional ones are exactly 60/n, which is what an authoring tool produces when it divides 60 by a target frame rate:
| Raw | 16.16 | = 60/n | fps |
|---|---|---|---|
0x00020000 |
2.00 | 60/30 | 30 |
0x00026666 |
2.40 | 60/25 | 25 |
0x00035555 |
3.33 | 60/18 | 18 |
0x0005745d |
5.45 | 60/11 | 11 |
0x00089249 |
8.57 | 60/7 | 7 |
0x00020000 — 30 fps — is by far the most common. Read as a plain integer it
is 131072, which yields a frame time of about 36 minutes.
Two different things both behave as animations:
ANIMfiles — an explicit animation chunk wrapping a sequence of cels, with a frame count and frame rate in the header.- Cel chains — a file that is simply several
CCBchunks concatenated. There is no header saying "this is an animation"; the only signal is that a secondCCBfollows the first. M2Suite detects this by walking the flat chunk chain and stopping as soon as it sees a secondCCB, which keeps the check cheap on large files.
Finding (verified): A PLUT may appear after the pixel data it
belongs to, not before. A parser that only applies a palette forward
produces a black or miscoloured frame. M2Suite back-fills a trailing PLUT
onto the preceding frame. Verified against StarBlade's boom.Anim.
Finding (verified): Real files append trailing data after the last
frame — padding, or another asset entirely. Treating a bad chunk size as
fatal throws away a perfectly good animation. The rule that works: if at
least one frame has already been decoded, stop cleanly at the bad chunk;
only error if nothing was decoded at all. Verified against Street Fighter's
PL05_ENDING.DAT (26 frames recovered) and syukyakuDEMO.dat (71 frames),
both of which previously failed outright with anim chunk '...' has bad size.
Finding (consistent): Some cel chains carry no header at all, so a synthetic one is derived from the first cel's CCB.
IMAG is a full-screen M1 image, typically 16-bit 0555.
Finding (verified): Some IMAG files declare a compression mode in the
header but store a complete uncompressed buffer anyway. Trusting the
header yields a mostly-black image. The fix is to check whether the payload
size already equals width × height × 2 and, if so, take it verbatim.
Verified against Yu Yu Hakusho's STOP.IMAG, which went from 99.7% black
pixels to a full image.
Both linear and LRForm (interleaved half-framebuffer) layouts occur, and the same de-interleave as for headerless cels applies.
A 3DO resource archive that wraps several assets in one file.
'RSRC' header
'RTBL' resource table
u32 count
records[count] 32 bytes each: type tag, id, offset, size, name
payload
Finding (verified): Files with an .AIFF extension are sometimes
actually RSRC bundles. Road Rash's Rash.AIFF is an RSRC containing 19
separate AIFF sounds, and a strict AIFF loader rejects it with
expected 'FORM', got 'RSRC'. M2Suite detects the RSRC magic at the top
of Aiff::load and unwraps to the first AIFF resource.
M2Suite writes these formats as well as reading them, for asset replacement and translation work. Two things are worth recording.
"Bitrate" is not a free parameter. Every 3DO codec is a fixed-ratio
scheme with no quality knob, so the bitrate falls out of
sampleRate × channels × bitsPerSample — 16, 8 or 4 bits. A UI that offers
a bitrate slider is lying about what the format can do; M2Suite reports the
resulting rate instead.
The delta codecs must be encoded against their own decoder. SDX2, SQS2 and CBD2 each carry per-channel history, so an encoder that disagrees with the decoder by one LSB does not sound slightly wrong — the error compounds and the file audibly drifts. The encoder here searches both the "exact" (even byte, history discarded) and "delta" (odd byte) forms for every sample and keeps whichever lands closer, which also stops error accumulating when a signal moves faster than the delta table can follow.
Measured round-trip error against a 440 Hz tone with deliberate transients:
| Codec | Ratio | RMS error (of 32000 full scale) |
|---|---|---|
| SDX2 | 2:1 | 41 (0.13%) |
| CBD2 | 2:1 | 33 (0.10%) |
| ADP4 | 4:1 | 1337 (4.2%) |
Container rule: AIFF has no compression tag, so anything other than PCM
needs AIFC — and an AIFC must carry an FVER chunk before its COMM.
3DO AIFF/AIFC files use several proprietary codecs identified by the
compression tag in the COMM chunk.
| Tag | Codec | Ratio | Notes |
|---|---|---|---|
NONE / (absent) |
Linear PCM 8 or 16-bit | 1:1 | |
SDX2 |
Square-Delta 2:1 | 2:1 | Squared-delta with sign; per-channel state |
SQS2 |
SDX2 variant | 2:1 | |
CBD2 |
Callisto Block Delta | 2:1 | |
ADP4 |
IMA/DVI ADPCM, 4-bit | 4:1 | Canonical 89-entry step table |
sowt |
Byte-swapped PCM | 1:1 | Not yet supported |
Finding (verified): AIFF sample rates are stored as an 80-bit IEEE 754 extended float, not an integer. It must be decoded properly; truncating the mantissa gives rates that are close but wrong, and the drift is audible over a long file.
Finding (verified): ADP4 is standard IMA ADPCM with the canonical
89-entry step-size table and the standard 16-entry index-adjust table — not
a 3DO-specific variant. Verified against Yu Yu Hakusho's SOUND/*.sc.
Technique (verified, and reusable): when you have a byte stream and several candidate interpretations, decode it under each and measure the lag-1 autocorrelation of the resulting samples. Real audio is strongly correlated sample-to-sample; a wrong interpretation is close to noise.
This settled the FILM audio question outright:
| Interpretation | Lag-1 autocorrelation |
|---|---|
Signed 8-bit PCM (what FFmpeg's segafilm assumes) |
0.013 |
| Unsigned 8-bit PCM | 0.014 |
| SDX2 | 0.923 |
No listening test needed — 0.92 versus 0.01 is not ambiguous.
3DO's streaming container, used for in-game movies. A flat sequence of chunks, each:
u32 tag
u32 size including this 8-byte header
u32 time stream-tick timestamp
u32 channel
payload
| Tag | Contents |
|---|---|
SHDR |
Stream header: tick rate, buffer sizing, channel map |
FILM |
Video subchunk (see below) |
SNDS |
Audio subchunk |
CTRL |
Control/marker |
FILL |
Padding |
FILM subchunks carry a further sub-tag: FHDR (film header), FRME
(keyframe), DFRM (difference frame).
Finding (verified): Files exist that begin with CTRL or FILL rather
than SHDR. Requiring a stream header at offset 0 rejects them. M2Suite
scans forward, counting payload chunks, and accepts the file if the chunk
chain is coherent. Verified against several Need For Speed Movies/*.Stream.
Finding (verified): DFRM difference frames must be decoded, not
skipped. A decoder that only handles FRME plays keyframes only and looks
like a slideshow — Coven's Intro.stream went from 242 usable frames to
6137 once DFRM was handled.
Finding (verified) — A/V sync: frame timestamps are in stream ticks, and
the tick rate in SHDR is not always usable. Deriving the effective rate
from the audio track's true decoded duration and then showing the newest
frame whose timestamp is due keeps long clips in sync, and tolerates
variable frame spacing. Assuming a constant frame rate drifts.
The same FILM payload without the DataStreamer wrapper — a QuickTime
derivative:
'FILM' size ...
'FDSC' film description: codec fourcc, width, height
'STAB' sample table: per-sample offset, size, timestamp, flags
'FRME' the samples themselves
Finding (verified) — the audio is SDX2, not PCM. The generic FILM/Sega
FILM description (and FFmpeg's segafilm demuxer) treats 8-bit FILM audio
as raw signed PCM. On 3DO discs it is SDX2-compressed at 22050 Hz, mono,
decoding to 16-bit. Decoding it as PCM produces loud noise. Established by
the autocorrelation test above against Yu Yu Hakusho's MOVIE/MV_01.film.
Finding (consistent): DFRM appears here too, alongside FRME, and
both must be accepted as sample types.
The video codec inside almost all 3DO FILM payloads (cvid).
Finding (verified): inter-frame (delta) decoding must follow the canonical algorithm exactly — in particular, the codebook update flags and the "skip" run semantics. An approximation produces progressive block corruption that only becomes obvious dozens of frames in. M2Suite's decoder was rewritten against the reference algorithm after exactly that failure.
Some discs carry MPEG-1 elementary streams inside 3DO containers.
Finding (verified): when an MPEG frame spans several chunks, the
continuation chunks use the tag FRM] and — unlike the first chunk —
carry no timestamp prefix. Stripping four bytes from every chunk
uniformly corrupts the stream. Verified against Oldsmobile's
72130.stream.
Finding (verified): M1VC containers nest. Pontiac's
ControlCenterUP.m1c wraps its real MPEG-1 640×480 payload six levels
deep. Unwrapping needs to recurse with a depth guard rather than peel one
layer.
The 3DO disc filesystem, present on M1 and M2 discs.
- 2048-byte logical blocks; the volume header is at block 0 and carries the
magic byte
0x01plus theZZZZZsynchronisation pattern. - Directories are block-chained; each entry has a type tag, block count, and a fixed-length name field.
- Filenames may be Shift-JIS on Japanese discs and must not be forced through a Latin-1 conversion.
Finding (verified): M2Suite's extractor was checked byte-identical against an established reference extractor across full discs, which is what makes disc extraction trustworthy enough to build everything else on.
.chd images are not supported — convert to .bin/.cue or .iso
first.
M2 runs PowerPC 602. Executables are ELF.
Finding (consistent): sections are frequently LZSS-compressed and must be inflated before disassembly, or the listing is noise. M2Suite detects and decompresses these automatically.
The disassembler resolves branch targets and applies symbol names from the ELF symbol table, and a second pass lifts the listing into ANSI-C-like pseudocode.
The 3DO ports of Alone in the Dark 1 and 2 keep the DOS engine's data formats, so everything here is little-endian. This is the most complete set of findings in the project.
u32 offsets[N] offsets[0] == 0; offsets[1] doubles as the table size,
so N = offsets[1] / 4 and entryCount = N - 1
per entry, at offsets[i]:
u32 skip size of the extra header, including this field
u8 extra[skip - 4]
u32 compressedSize
u32 uncompressedSize
u8 compressionType
u8 compressionFlags
u16 padding
u8 payload[] at entryOffset + 16 + extraLen + padding
Finding (verified) — the compression is PKWARE DCL "implode", not LZSS.
This cost real time. compressionType == 1 was initially assumed to be the
LZSS used elsewhere in the engine, and that assumption was empirically
disproved before the right answer was found: the first control byte 0x0F
decodes as "four literals, then a back-reference to offset 789", which is
impossible when only four bytes of output exist. The actual codec is PKWARE
Data Compression Library implode, decoded with Mark Adler's public-domain
blast. In the compression flags, bit 2 (0x04) means a literal Huffman
tree is present and bit 1 (0x02) selects the 8 KiB dictionary.
compressionType == 0 is stored; 4 is a deflate-like variant.
Per the fitd reference implementation (hqr.cpp, createBodyFromPtr):
u16 flags
s16 bbox[6] ZVX1, ZVX2, ZVY1, ZVY2, ZVZ1, ZVZ2
u16 scratchSize
u8 scratch[scratchSize]
u16 numVertices
s16 vertices[numVertices][3]
-- if flags & INFO_ANIM (2):
u16 numGroups
u16 groupOrder[numGroups]
u8 groupRecords[numGroups][recSize]
recSize = 0x18 if flags & INFO_OPTIMISE (8), else 0x10
u16 numPrimitives
primitives[numPrimitives]
Flags: INFO_ANIM = 2, INFO_TORTUE = 4, INFO_OPTIMISE = 8.
Finding (verified) — the field at offset 14 is scratchBufferSize, not
the vertex count. An early heuristic that searched for a vertex count near
the header found a value that "fit" and produced geometry that looked
almost right. It is actually the size of a variable-length scratch buffer
that sits before the real count, so every model was being read at the
wrong offset. There is no shortcut here: the layout must be walked exactly.
Finding (verified) — the single most consequential AITD finding so far.
A body with INFO_ANIM stores its vertices in group-local space. Each
group (bone) owns a run of vertices, and every one of them is an offset
from the position of that group's base vertex — the vertex it hangs off in
its parent. A parser that stops after reading the vertex array gets a mesh
whose limbs are all bunched around the origin.
To resolve it, walk the groups in stored order, in place:
for each group g, in the order they appear in the file:
base = vertices[g.baseVertex] // already in model space
for k in 0 .. g.vertexCount-1:
vertices[g.start + k] += base
Two details are load-bearing:
- Stored order matters. Groups are laid out parents-first, so by the time a child is reached its base vertex has already been moved into model space and the offsets cascade correctly down the chain.
- It must be in place. Resolving against a pristine copy of the vertex array breaks every group below the first level, because children need the parent's resolved position, not its original one.
start and baseVertex are stored as byte offsets — divide by 6.
Group record layout, following the group-order table (one u16 per group):
| Field | Size | Notes |
|---|---|---|
start |
u16 | first vertex, byte offset (÷6) |
numVertices |
u16 | how many vertices the group owns |
baseVertices |
u16 | base vertex, byte offset (÷6) |
orgGroup |
s8 | parent group index, −1 at the root |
numGroup |
s8 | |
state.type |
s16 | 0 = rotate, 1 = translate, 2 = zoom |
state.delta[3] |
s16 × 3 | animation deltas, zero in the bind pose |
state.rotateDelta[3] + pad |
s16 × 4 | AITD2 only (INFO_OPTIMISE) |
So the record is 0x10 bytes on AITD1 and 0x18 on AITD2. Reading the wrong stride swallows the primitive count and the whole primitive list decodes as garbage.
How it was caught: LISTBOD2.PAK entry 12 (AITD1) is Emily Hartwood.
It rendered as a jumble of overlapping shards. Comparing against a
reference render of the same body made it obvious the geometry was present
but misplaced, not corrupt — which pointed at a transform rather than a
parse error. Pinned by regression tests
(test_aitd_synth.cpp, testGroupHierarchyResolves).
u8 type
| Type | Name | Payload after type |
|---|---|---|
| 0 | Line | u8 subType, u8 color, u8 even, u16 points[2] |
| 1 | Poly | u8 n, u8 subType, u8 color, u16 points[n] |
| 2 | Point | u8 subType, u8 color, u8 even, u16 points[1] |
| 3 | Sphere | u8 subType, u8 color, u8 even, u16 size, u16 points[1] |
| 4 | Disk | (not observed in 3DO data) |
| 5 | Cylinder | (not observed in 3DO data) |
| 6 | BigPoint | as Point |
| 7 | Zixel | as Point |
| 8 | PolyTexture8 | as Poly |
| 9 | PolyTexture9 | as Poly, then u8 uv[n][2] |
| 10 | PolyTexture10 | as Poly, then u8 uv[n][2] |
Finding (verified) — vertex indices are stored as byte offsets. Divide
each u16 by 6 (three s16 per vertex) to get the vertex index.
Finding (verified) — PolyTexture 9 and 10 append a (u,v) byte pair per
point; type 8 does not. These UVs sit inline in the primitive stream, so a
parser that skips them desynchronises every following primitive. The
symptom is distinctive: models grow long thin spikes shooting off into
space, because garbage indices reference distant vertices. This is pinned by
a regression test (tests/libm2model_tests/test_aitd_synth.cpp).
Finding (verified) — there is no textured geometry in the 3DO builds. A sweep of every PAK in both games found zero primitives of type 8, 9 or 10 in any real model archive:
| Archive | Game | Bodies | Primitive types present |
|---|---|---|---|
LISTBODY.PAK |
AITD 1 | 272 | 0, 1, 2, 3, 6, 7 |
LISTBOD2.PAK |
AITD 1 | 272 | 0, 1, 2, 3, 6, 7 |
4LSTBODY.PAK |
AITD 2 | 551 / 553 | 0, 1, 2, 3, 6, 7 |
Every polygon is flat-shaded from the palette. So "sample the PolyTexture
primitives" has no data to act on for these two ports — the UV parsing is
implemented and tested so that a build which does use them (the PC
releases, or AITD 3) decodes correctly, but on 3DO there is nothing to
sample. The reference implementation fitd reaches the same conclusion from
the other direction: it renders PolyTexture primitives with
processPrim_Poly, i.e. flat, and never samples a texture.
Determined by sweeping every .PAK and counting entries that parse as real
geometry (≥4 vertices and ≥2 primitives):
| Archive | Contents |
|---|---|
LISTBODY.PAK, LISTBOD2.PAK, *LSTBODY.PAK |
3D models |
LISTANIM.PAK, LISTANI2.PAK, 5LSTANIM.PAK |
Model animations |
ETAGE00..15.PAK |
Floors / rooms |
MASK00..15.PAK |
Camera masks (2D clipping polygons) |
ANIM00..15.PAK |
Room animations |
ITD_RESS.PAK, JAP_RESS.PAK |
Engine resources |
LISTSAMP.PAK, chsamp.pak |
Sound samples |
0LSTMAT / 1LSTLIFE / 2LSTTRAK / 3LSTHYB |
Scripts and tables |
Trap: almost any blob can be coerced through the body parser and yield one or two "valid" bodies. MASK PAKs in particular produce a handful of false positives. Identifying a model archive therefore requires a majority of sampled entries to carry real geometry, not merely one.
The room backgrounds. Three container extensions, one payload layout:
u16 width 320 on every file seen
u16 height 200
u8 palette[768] 256 RGB triples, 8 bits per channel
u8 pixels[w*h] one palette index per pixel
.bob— a single un-padded page.pad— a single padded page.pics— one or more pages, each padded
Finding (verified) — by exact arithmetic. 4 + 768 + 320×200 = 64772,
which is the exact byte size of CAM8003.BOB. That is not a coincidence,
and it confirms the palette size and pixel depth in one step. Decoded output
was then visually checked against recognisable rooms.
Finding (verified) — pages pad to a 4 KiB boundary, and dimensions are not fixed. This one was originally got wrong in an instructive way. Every AITD2 backdrop first sampled was 320×200, whose payload of 64772 bytes rounds to exactly 65536 — so "pages are padded to 64 KiB" fitted perfectly and was hard-coded. It is wrong:
| File | Size | Payload | Stride | Pages |
|---|---|---|---|---|
camera00.pics (AITD1) |
240×200 | 48772 | 49152 | 5 |
camera01.pics (AITD2) |
320×250 | 80772 | 81920 | 15 |
camera03.pics (AITD2) |
320×200 | 64772 | 65536 | 2 |
The rule is stride = roundUp(4 + 768 + w×h, 4096), and every page in a
file shares the first page's dimensions. Under the 64 KiB assumption both
non-320×200 files were rejected outright — a reminder that a constant which
fits every sample can still be a coincidence, and that the cheapest defence
is a sample whose dimensions differ.
Finding (verified). AITD2 ships its sound effects as many complete
FORM/AIFF files concatenated into one container, each starting on a
2048-byte boundary (one CD sector) with the gaps zero-padded.
LISTSAMP.CAT holds 193 sounds; .FRE and .jpn carry the localised
sets.
The sector alignment is the reliable signal. Scanning for the FORM tag at
any offset also matches the byte pattern occurring inside sample data —
three false positives in LISTSAMP.CAT alone — so only sector boundaries
are considered. Walking the chain sequentially does not work either: the
file has gaps that are not sounds, and a sequential walk stops after 66 of
the 193.
A catalogue is told apart from a single sound structurally: it opens with a
FORM whose declared size leaves most of the file unaccounted for.
Finding (verified). AITD's visuals are the pre-rendered backdrops, so a "room" stores no mesh. What it stores is a set of axis-aligned boxes: colliders the player walks on and bumps into, and triggers that fire scripts. Reassembled, they give a floor's true layout and scale.
Every room of a floor lives in the archive's first entry, behind a u32
offset table; the second entry is camera data. (A later engine revision put
one room per entry — AITD-roomviewer switches on the entry count. Both 3DO
builds use the table form, with exactly two entries per archive.)
entry 0:
u32 roomOffsets[] offsets[0] is BOTH the table size and the first
room's offset — the rooms begin immediately after
the table. A zero or out-of-range slot ends the list.
per room, relative to its offset:
u16 colliderOffset @ 0 relative to the ROOM, not the entry
u16 triggerOffset @ 2
s16 position[3] @ 4 room origin; multiply by 10 for engine units
u16 cameraCount @ 10
u16 cameraIds[n] @ 12
at colliderOffset (and likewise at triggerOffset):
u16 count
count x 16 bytes:
s16 lowerX, upperX, lowerY, upperY, lowerZ, upperZ
s16 id @ 12
u16 flags @ 14
Collider flags: 0x02 underground floor, 0x04 link to another room,
0x08 interactive.
Two traps. The box-list offsets are relative to the room, not to the entry — using entry-relative offsets yields zero rooms rather than wrong ones, which at least fails loudly. And the room position is in units ten times coarser than the box coordinates, so a floor assembled without the ×10 collapses into a heap.
Verified against both 3DO games; AITD1's eight floors yield 1–14 rooms each (ETAGE03: 14 rooms, 178 colliders, 43 triggers, 72 cameras) and render as a recognisable Derceto floor plan.
Derived from tigrouind's AITD-roomviewer
(RoomLoader.cs).
The 256-colour AITD palette is a fixed table, transcribed from
AITD_PakEdit's AloneFile::palette and vendored as
libs/libm2model/src/AitdPalette.cpp. Primitive color fields index it
directly.
Finding (verified): AITD is Y-down — negative Y is up. A standing
character has feet near y = 0 and head near y = -1700. Because screen Y
also increases downward, model Y maps straight to screen Y and models render
upright with no flip. Exporters targeting Y-up tools (OBJ) must negate Y
and Z to preserve handedness; M2Suite's OBJ exporter does.
AITD_PakEdit ships hand-curated JSON databases (AITD1_CD_PAK_DB.json,
AITD1_floppy_PAK_DB.json, …) that name the contents of each archive
entry. They are the community's accumulated identification work and turn
"Model 12" into "Emily Hartwood".
{ "all_PAKs": {
"CAMERA02.PAK": {
"25": { "default_compr": 1, "info": "1st Floor Library (Secret Room)", "type": 2 }
} } }Keys are archive filenames, then entry index as a string. info is "?"
when the entry has not been identified. The type field is a content
class:
type |
Meaning |
|---|---|
| 0 | Unclassified / other |
| 1 | Text (in-game messages, readable documents) |
| 2 | Background image |
| 3 | Floor / room layout |
| 4 | Camera data |
| 6 | Full-screen image sequence |
| 7 | 3D model (body) |
These databases doubled as an independent check on the renderer. After
the bone-group fix, a contact sheet of LISTBOD2.PAK was rendered blind and
then compared against the database names: entry 12 "Emily", entry 5 "4 Pane
Window", entry 0 "Wooden Chest (Closed), Loft", entry 20 "Indian Cover" —
every one matched what had been drawn. Cheap, and much stronger evidence
than eyeballing a single model.
M2Suite reads any *PAK_DB.json sitting beside the archive or one level up
and uses it to label entries. It does not bundle one — the databases
belong to their authors, and are game-specific.
Two bugs here produced the "broken models" symptom, and neither was a parsing problem:
The declared bounding box is not the geometry. It doubles as a collision volume and can be an order of magnitude larger than the mesh. Framing the camera on it renders the model as a speck in the corner. Fit to the actual vertex extents.
Unreferenced vertices drag the centre. The vertex array contains vertices no primitive references — bone roots and animation helpers, typically sitting at the origin. Including them in the extents pushes models low in the frame and clips them off the bottom edge. Fit only to vertices that a primitive actually uses.
The combination that works: compute the centre from the extents of referenced vertices, then scale by the bounding sphere radius rather than per-axis extents. The sphere radius is rotation-invariant, so the model keeps a constant size and never clips as the camera orbits. Both properties are pinned by regression tests.
Painter's algorithm is not enough. Depth-sorting whole faces renders interpenetrating geometry incorrectly — limbs crossing a torso visibly tear. A per-pixel depth buffer, with depth interpolated across each scanline span, fixes it. Sorting is kept as a cheap first pass.
AITD has no lighting model — do not add one. The artists baked shading into their choice of palette index per face, which is why adjacent faces already read as lit. Applying a light term on top double-shades the model; combined with a two-sided winding test it dimmed most of the mesh, and Emily Hartwood came out near-black next to a reference render of the same body. The faithful mode renders the palette colour unmodified. A neutral grey mode keeps the shading, because there the point is to read silhouette and form rather than to reproduce the original.
Finding (verified). Both ReadySoft 3DO titles — Dragon's Lair and Space
Ace — ship their full-motion video in one container format, keyed by the
BNDY ("boundary") magic:
repeat:
u32 blockSize covers the whole block from this word
u32 flags 0xF0000000 on Space Ace, 0 on Dragon's Lair
'BNDY'|'ANDY' u16 index u16 mark u8 pad[16] 24-byte block header
'FORM' u32 size 'SASQ' <chunks> a standard EA-IFF-85 FORM (Space Ace)
blockSize chains the blocks; the first block of a movie is tagged BNDY
and every continuation is ANDY. Verified: every one of the 29 Dragon's
Lair .DAT files and 21 Space Ace .DAT files parses as such a chain.
Chunks inside the FORM are 4-byte aligned (not 2 — the 2-byte
assumption silently loses every chunk after the first odd-sized one):
| Chunk | Payload |
|---|---|
BPAL |
big-endian RGB555 palette entries, from index 0 |
BAC8 |
keyframe: s16 x, s16 y, u16 w, u16 h then w*h raw 8-bit indices |
MVB8 |
identical layout — a raw sub-rect blit, used for scroll-in strips |
MOVE |
s16 x, s16 y — the viewport origin on the canvas |
CELL |
same rect header, then a standard 3DO packed coded 8-bit cel |
COPY |
zero-length marker |
FSND / ASND |
raw signed 8-bit PCM, 22050 Hz mono |
The picture model is a palettised canvas that can be larger than the
288×216 viewport: BAC8/MVB8 supply raw pixels, MOVE pans the view, and
each frame's CELL composites a packed cel over whatever is already there
(transparent runs leave the canvas untouched). Frames are therefore deltas
and must be decoded in order.
The CELL payload is the stock 3DO packed-cel encoding, confirmed by
arithmetic rather than assumption: a fully transparent row reads
00 00 BF BF BF BF 9F 00, i.e. a row-offset word of 0 then transparent runs
of 64+64+64+64+33 = 288 pixels — exactly the frame width — followed by
end-of-row, in (0 + 2) * 4 = 8 bytes. Each row is u16 offset (row length
in 32-bit words, minus two) then control bytes whose top two bits select
0 end-of-row, 1 literal run, 2 transparent run, 3 repeat, with the
low six bits holding count−1.
Timing. Every ASND is exactly 1838 bytes = 1838/22050 s, so the movies
run at 11.996 fps — 12 fps. Frame 0 carries an FSND preload that can
exceed a second, so M2Suite paces frames off the running audio-byte total
rather than a flat 12 fps (a flat rate drifts ~1.2 s over a 48 s movie).
The audio is raw signed 8-bit PCM rather than SDX2 — decided by measurement: as PCM the track has 0.83 lag-1 autocorrelation and never clips; decoded as SDX2 it falls to 0.79 and saturates both rails.
BAC8 keyframes are interlaced. Even and odd scanlines are two
temporally distinct fields woven together, so fast-motion keyframes comb.
That is verified, not assumed: within a single file the woven-versus-
deinterlaced roughness ratio ranges from 1.1 on near-static frames to 11.0
on fast motion. A structural row-order error would be constant; content
dependence means real motion. Weaving reproduces what the format stores and
what the 3DO fed its CRT. Two alternative row orders were tested and
rejected — the stacked-half reading leaves a seam 4–5× rougher than the
surrounding image, and the fields are not a spatial offset of one another.
Dragon's Lair's blocks carry flags = 0 and no IFF FORM: the payload
begins immediately after the 24-byte block header. The SASQ chunk decoder
does not apply, so M2Suite maps the block chain and reports it rather than
drawing guessed pixels.
ONE.3DO (Dragon's Lair) is not a BNDY container — a small 5 KB file,
likely an index or config — and is left as a hex view.
Finding (verified). Paul.pict (3DO Gamepack) is a genuine Macintosh
PICT: a 512-byte all-zero header followed by a QuickDraw picture. The
version is read from the record after the header — 0x0011 0x02FF marks
PICT v2 (extended QuickDraw); Paul.pict is v2, 240×240. M2Suite identifies
it and reports the version and bounding rect; where Qt's image plugins can
read the PICT it is shown, but the QuickDraw opcode stream (bitmaps,
regions, text) is not decoded here.
Finding (verified). Some game .aif files are raw samples with no
FORM wrapper (3DO Gamepack onside/fb_au/003.aif), which made a strict
IFF parser throw. The loader now falls back to treating a non-FORM,
non-RSRC input as raw signed 8-bit PCM, mono, 22050 Hz so it plays
rather than erroring — a best-effort interpretation of a stripped file, not
a claim about a stored header. (Measured autocorrelation of 003.aif is
weak under both 8- and 16-bit readings, ~0.22 vs ~0.15, so it may be
delta-coded; 8-bit is the better default.)
Big-endian throughout, verified against the real 331 MB GXdata/bigfile:
u32 entryCount
u32 reserved[3] zero in every container seen
entryCount x {
u32 id filename hash — no names are stored
u32 size
u32 offset 2048-byte (disc sector) aligned
}
payload
Each entry's offset is the previous entry's end rounded up to 2048, which is what confirms the field order.
Finding (verified) — the format is recursive, and nothing is compressed. This was the whole puzzle. Entries are frequently complete bigfiles themselves, using the identical header:
| Top-level entries | 244 |
| …of which are containers | 158 |
| Nested entries inside them | 7,991 |
| Leaves after recursing | 8,077 (97.4% of the file) |
| Nesting depth | 2 |
Flat extraction stops one level short and leaves most of the game's assets sealed inside blobs that look like compressed data because they are opaque. They are not. Every leaf was measured for byte entropy and byte-to-byte correlation and not one is compressed — no entry has the high-entropy, low-correlation signature that compression leaves.
So there is no decompressor to write. The answer to "how do I decompress the bigfile" is that you don't; you recurse.
Detection must be strict. The container check requires the three reserved words to be zero, every offset to be sector-aligned and inside the blob, and the records to account for all but the trailing pad sector. Without the zero-reserved-words guard the parser "finds" containers inside ordinary sample data, and a false positive shreds a real asset into noise — worse than leaving it whole.
Finding (verified by byte comparison.) Fifteen top-level entries are each exactly 12,804,096 bytes, at fifteen distinct offsets, with fifteen distinct name hashes — and all fifteen are byte-identical. That is 171 MB of a 316 MB archive holding one repeated blob.
This is not a bug in the reader: the table has 244 distinct offsets and no two entries share one. It is the classic CD-ROM seek-time optimisation — place a copy of the shared asset bank near each level's data so a 2x drive never has to seek across the disc. Worth knowing before anyone concludes their extraction is duplicating output.
The release bigfile stores only name hashes, which is why extracted
entries are numbered. The Gex alpha (3DO Gamepack GEX_alpha) solves
this: instead of one monolithic hash-keyed archive, it keeps a small per-
level .big index for each named level directory (kungfu1.dir,
jungle2.dir, grave1.dir, …), and those indexes store the plaintext
asset names next to their hashes:
u32 count
u32 reserved[3]
count x { u32 hash; u32 field1; u32 field2 }
<name table> NUL-separated paths, e.g. "kungfu1.dir/map.map",
"paralli/carton1b.par"
Finding (verified). The hashes in the alpha .big files match entries
in the release bigfile's nested tree — e.g. 2A39723E is
paralli/carton1b.par, and it is a real leaf in the release archive. Cross-
referencing the alpha's hash→name map against the release bigfile at every
depth names 41 release entries with meaningful paths: level maps
(grave1.dir/map.map, cartoon5.dir/map.map), parallax backgrounds
(paralli/jungle2b.par), and sound banks. This is the same win as the AITD
name database, but derived rather than transcribed.
Consistent (partial). The name hash itself is a 4-byte XOR fold —
adjacent level names differ in exactly the hash bytes their changed
characters XOR (glue1→glue3 flips only the affected bytes). A brute force
over fold parameters reproduces 34 of 77 alpha hashes, so the family is
confirmed but the exact permutation is not yet pinned; the direct hash→name
table from the alpha is what does the naming today. Fully reversing the fold
would name all ~8,000 release leaves, not just the 41 the alpha references.
Each copy of that blob is a container of 516 entries whose ids are a counter (1…0x204) rather than hashes — 442 of them the same size. They are sounds:
s32 unknown varies, always negative (-1 .. -35248)
u32 length payload bytes
s32 loopStart -1 when the sound does not loop
s32 loopEnd -1
u8 pcm[length] signed 8-bit
(length + 16) rounded up to 4 equals the entry size for 468 of 516
entries — the remainder are the bank's non-sound assets, and that identity
is what tells the two apart. 492 of 516 carry -1/-1 loop points. At
22050 Hz the common 23,499-byte payload is ~1.07 s, and the amplitude
envelope varies the way speech does rather than holding steady like a test
tone.
The sample rate is not stored, and it is not uniform. 22050 Hz is a
default chosen because it makes typical entries about a second long, but a
subset of sounds audibly play too fast at that rate — they were authored
lower. Nothing in the per-sound header encodes the rate: field[0] looked
like a candidate but its high 16 bits are always 0xFFFF and its value
tracks the content (identical sounds share it), not a rate. M2Suite
therefore exposes the playback rate as a control rather than guessing.
The whole tree is 59% redundant. The 8,077 extracted leaves hold only
729 distinct contents. This is the fifteen-fold level-bank duplication
seen from the other side: 441 distinct sounds each appear 15 times (once per
level bank), and one appears 750 times. So "the same sound under
different names in different folders" is exactly that — byte-identical
copies, a CD seek-time optimisation, not variants. A recurring field[0]
value is just a recurring sound.
M2Suite's extractor therefore de-duplicates by content hash and sorts leaves
into type folders (audio/, images/, video/, geometry/, data/)
rather than mirroring the nameless container nesting. On Gex that turns
8,077 files (308 MB) into 729 (126 MB): 415 audio, ~312 data/geometry, 2
images, plus the video kept whole. The on-disc layout was optimised for a 2x
CD drive; the extracted layout is optimised for a person browsing it.
Recognition, measured through M2Suite's own sniffer over the recursively extracted tree: 7,020 of 8,077 leaves (86.9%) are sounds, 124 minutes of audio. As a falsification check the decoded output was measured for sample-to-sample correlation — 99.4% score above 0.5, mean 0.831 — so the detector is not simply claiming everything.
Each movie is four files sharing a stem:
| File | Contents |
|---|---|
.DUK |
video data, frames laid end to end |
.FRM |
frame index — big-endian u32 offsets into the .DUK |
.HDR |
48-byte movie header |
.TBL |
4096 bytes = 256 entries of 16 (codebook) |
Verified: the frame index is strictly ascending and covers the video
data in all four movies (614, 917, 423 and 209 frames). Each frame begins
with a 12-byte preamble whose bytes 8–11 are the constant magic
F7 7F 05 E4, which is what makes a .DUK identifiable on sight.
The .HDR is three big-endian u32 fields, a u16 pair, then eight signed
(x,y) motion-vector pairs — the same eight across all four movies:
(0,-2) (2,-6) (6,-12) (12,-12) (0,-1) (1,-2) (3,-4) (5,-4).
Consistent (hypothesis) — the sidecar files externalise the codec's
tables. This is the promising lead, and it is why FFmpeg's truemotion1
cannot be pointed at these frames directly. Stock TrueMotion 1 selects,
per frame, one of a set of built-in tables via two header indices:
deltaset (a predictor/delta set) and vectable (a vector codebook). The
Gex sidecars look like exactly those tables lifted out of the binary:
.HDRends in eight signed (x,y) pairs — the same eight across all four movies — where TrueMotion 1 keeps its predictor/delta vectors..TBLis 256 × 16 bytes, the shape of a per-code vector codebook.
This is graded consistent, not verified: the .TBL layout does not yet
match TrueMotion 1's variable-length sel_vector_table byte-for-byte
(FFmpeg builds that table as, per code, a length byte then len/2 delta
pairs), so either the .TBL is a pre-expanded form or the mapping is not
one-to-one. Confirming it is the next step.
Verified — the frame body is compressed, not raw. Frame 0's body was
rendered as 8-bit grayscale and 16-bit 0555 RGB at every plausible width
(64…320); every result is noise. So the data is genuinely vector-quantised
— there is no width at which it is a raw bitmap, and the codec must be
implemented. Frame data measures ~6.9 bits of entropy per byte, consistent
with VQ.
Verified — each frame carries two data regions. The 16-byte preamble is
two big-endian length words, the magic, then a u32 (zero on frame 0):
u32 region0_len e.g. 1518 on BASK frame 0
u32 region1_len e.g. 6644
u8 magic[4] F7 7F 05 E4
u32 (zero on the first frame)
Decoded (verified). The two regions are audio then video, and the record layout is:
u32 audio_size
u32 video_size
u8 audio[audio_size]
u8 video[video_size]
u8 sector_padding[]
The decoder is in libm2duck, ported from The Ur-Quan Masters 0.8.0
(src/libs/video/dukvid.c and src/libs/sound/decoders/dukaud.c), which
carries a working decoder for exactly this 3DO variant. Because that source
is GPL-2.0-or-later, libm2duck is GPL-2.0-or-later — compatible with
M2Suite's GPL-3.0, and recorded in
THIRD_PARTY_LICENSES.md.
Header (.HDR, 48 bytes).
u32 version 3 selects the version-three decode loop
u32 screen_x_offset metadata for the game's own compositor
u32 screen_y_offset
u16 width_in_4px_blocks
u16 height_in_4px_blocks
s16 luma_prototypes[8]
s16 chroma_prototypes[8]
Codebook (.TBL, 4096 bytes). 256 vectors of 16 bytes. Byte zero of
each vector is the count of meaningful following items; the rest select
pairs from the .HDR prototypes. At load time these expand into two signed
256 x 16 delta tables (luma and chroma). The low bit of the last delta is
an end-of-sequence marker telling the frame loop to fetch another vector
byte. This table is precisely what stock TrueMotion 1 keeps compiled in,
which is why FFmpeg and ScummVM cannot decode these payloads directly.
Video subframe. A 16-byte 3DO prefix, then vector-coded picture data
from offset 0x10. The first big-endian 16-bit word selects the loop:
0x0300 picks the version-three path, anything else the original. Decoding
works in four-pixel-wide blocks, reconstructing vertically paired RGB555
pixels — each 32-bit accumulator holds the upper scanline's pixel in bits
31–16 and the lower scanline's in bits 15–0.
Audio subframe.
u16 magic 0xF77F
u16 numsamples encoded bytes; two decoded nibbles per byte
u16 tag
u16 index_left
u16 index_right
u8 adpcm[numsamples]
22050 Hz stereo, signed 16-bit after decoding. Nibbles alternate left/right (high nibble first). Predictor state persists across frames while the step indices are re-seeded by every subframe, and the step table differs from canonical IMA ADPCM — the 3DO/UQM values are required.
Timing. A constant 14.622 fps; audio is the sync authority.
| Movie | Version | Size | Frames | Duration |
|---|---|---|---|---|
| BASK | 1 | 280x152 | 614 | 41.99 s |
| GEXIALL | 1 | 280x176 | 917 | 62.71 s |
| GEXOALL | 1 | 280x176 | 423 | 28.93 s |
| LOGOTHIN | 3 | 280x184 | 209 | 14.29 s |
Verified against all four movies: BASK decodes a clean basketball scene, LOGOTHIN (the version-three path) the Crystal Dynamics logo, and every decoded audio track's duration matches the table above.
Reference sources:
- The Ur-Quan Masters 0.8.0 —
dukvid.c/dukaud.c, the working decoder this port is derived from. - FFmpeg
libavcodec/truemotion1.c— the canonical open TrueMotion 1 decoder;truemotion1_decode_headerandgen_vector_table16show howdeltaset/vectabledrive the built-in tables that Gex externalises. - The Duck Corporation's own TrueMotion "S" SDK, including the
tm1.0/decoder source and a Sega Saturn port, both in theDuckTruemotionSreference tree.
Of the 729 distinct leaf contents, 415 are sounds and 314 are not. The non-sound blobs are the level-geometry and sprite candidates, and they carry no magic. Characterised by content:
| Kind | Distinct | Notes |
|---|---|---|
| Structured data | 288 | The bulk. Runs of 0xFF, then mixed index bytes and big-endian floats (46 C0 00 00 = 24576.0) — level geometry and object data. |
| Sparse tables / tilemaps | 14 | Mostly zero, small |
| High-entropy (packed pixels?) | 7 | Larger, ~7.5+ bits/byte |
| 3DO IMAG images | 2 | Already decoded elsewhere |
| Offset table / index | 1 | The very first leaf |
These are individually reachable, listed, and shown with a hex header in the app, but their internal formats are not decoded — genuine reverse-engineering targets rather than transcription.
| Path | Contents |
|---|---|
LaunchMe, gxduckPlay, gxextra, gxsupprt |
3DO M1 ARM executables |
System/ |
42 more ARM binaries plus 65 3INS DSP instruments |
layout.json |
written by the disc extractor, not game data |
The ARM executables all begin E1 A0 00 00 — MOV r0,r0, a no-op, as the
entry point. 3DO stores ARM code big-endian, so that byte is byte 0.
M2Suite's ARM detector originally tested byte 3 (little-endian ARM habit)
and so never fired on a real 3DO executable; opening these files is what
exposed it.
Documented so nobody repeats the search:
- Gex leaf assets — the container nesting is solved and every leaf is now reachable and uncompressed, but most leaf payloads are still game-internal formats with no magic. The sound bank above is decoded; the level geometry and sprite data are not.
- PEBM (Oldsmobile) — the container parses, but it stores no width or height anywhere, so the pixel buffer cannot be laid out. Solving this needs either a companion file that carries the dimensions or a sample whose dimensions are known from elsewhere.
Corrections and additions are welcome — see CONTRIBUTING.md. A finding backed by a reproducible observation is worth more than a plausible theory, and a finding that disproves something on this page is worth most of all.