Skip to content

Latest commit

 

History

History
1303 lines (1023 loc) · 53.8 KB

File metadata and controls

1303 lines (1023 loc) · 53.8 KB

Format findings

Everything M2Suite knows about 3DO and Panasonic M2 data, written down so the knowledge survives independently of the code. Each entry says what the layout is, how it was established, and — where it matters — which published description is wrong.

Conventions: u8/u16/u32 unsigned, s16 signed, [n] array. 3DO M1 and M2 chunk containers are big-endian; Alone in the Dark data is little-endian (it is a DOS port). Byte offsets are from the start of the structure being described.

Contents


Reading this document

Findings are graded, because not all of them are equally certain:

  • Verified — decoded output was checked against something independent (a reference implementation, a known-good render, a statistical test, or an exact file-size match).
  • Consistent — the layout explains every file we have, but nothing external corroborates it.
  • Assumed — a working guess that has not failed yet.

IFF-style containers

3DO and M2 both use FORM-style chunk containers, but with different padding rules, and the difference is not cosmetic.

u32 tag         four-character chunk id
u32 size        chunk size INCLUDING these 8 header bytes
u8  payload[size - 8]
u8  pad[]       to the alignment boundary

Finding (verified): M2 chunks align to 4 bytes; classic IFF-85 chunks align to 2. A parser that assumes one alignment silently desynchronises on the other after the first odd-sized chunk. libm2core's IffForm accepts both and picks per container.

Container ids seen in the wild: FORM, CAT , LIST. A CAT behaves as a concatenation of FORMs and must be walked, not treated as a leaf.


Textures — UTF / M2TX

M2's native texture format. FORM+TXTR, or CAT +TXTR for a bundle of several textures in one file.

Sub-chunks:

Chunk Contents
M2TX Header: dimensions, texel format, LOD count, flags
M2PI PIP — the colour lookup table (palette)
M2CI DCI — colour/alpha component description
M2TD The texel data itself, LOD by LOD
M2LR LOD run lengths for compressed texel data

Finding (consistent): LOD 0 is the largest mip level and appears first. Texel data for later LODs follows contiguously; the per-LOD sizes come from the header's dimensions, halved and floored per level, not from a stored table.

Currently decoded: PIP-indexed, uncompressed, 8-bits-per-index. Other texel formats (RLE-compressed runs, direct/non-palette formats) are detected and rejected rather than guessed at — see LIMITATIONS.md.

Ground truth for this format is the SDK Mercury texture library C source; see REFERENCES.md.


OFST-wrapped cels

Finding (verified). Some 3DO Gamepack art wraps a normal cel in an OFST offset-table header rather than starting with CCB directly:

'OFST'  u32 celOffset  u32 sibling[...]
        ...            the real 'CCB ' chunk begins at celOffset

celOffset points straight at the CCB; the remaining u32s index sibling resources. Skipping to the CCB decodes the image normally — verified against starC.cel (40×41), ring3.cel (320×240), shadow.cel (16×16) and AJLbanner.cel (a 320×240 title banner). The loader unwraps OFST transparently.

Cels — CCB / PDAT / PLUT

The 3DO M1 sprite/bitmap primitive. A cel is a chunk chain:

Chunk Contents
CCB Cel control block: flags, dimensions, bit depth, pixel-data offsets
PDAT The pixel data
PLUT Palette lookup table (16-bit 0RRRRRGGGGGBBBBB entries)
XTRA Optional extra description

Bit depths 1, 2, 4, 6, 8 and 16 all occur. Depths below 8 are coded (palette-indexed via the PLUT); 16-bit is direct colour.

The pixel processor (PIXC / PPMPC) — why coded cels look flat without it

Finding (verified). An 8bpp coded pixel is not simply a palette index. The low 5 bits index the 32-entry PLUT; the top 3 bits are the AMV, a per-pixel brightness multiplier that the cel engine's pixel processor applies as it draws. Masking them off — the obvious reading — throws away all of a sprite's shading and renders it posterised.

The CCB's PIXC word holds two 16-bit PPMPC values: PPMP_0 in the low half for pixels whose SSB is clear, PPMP_1 in the high half for pixels with it set. For a coded cel the SSB is bit 15 of the PLUT entry. Field layout, from 3DO Portfolio 2.5 Interfaces/2p5/includes/hardware.h:

Bits Field Meaning
15 1S first source: 0 = cel pixel, 1 = framebuffer
14–13 MS multiplier source: 0 = the CCB's MF, 1 = the pixel's AMV (PPMPC_MS_PIN)
12–10 MF multiply factor, stored as value−1 (0b111 = ×8)
9–8 SF divide factor — 0b00 = 16, 01 = 2, 10 = 4, 11 = 8
7–6 2S second source: 0 = none, 1 = the CCB's AV, 2 = framebuffer
5–1 AV add value
0 2D second-source doubling

Two traps worth naming. PPMPC_SF_16 is 0x0000, so a zero SF field means divide by sixteen, not by one. And the AMV is encoded the same way MF is — as value−1 — so a fully lit pixel (AMV 7) multiplies by 8 and, with the usual divide by 8, reproduces the palette colour exactly. Reading the AMV as a raw 0–7 caps every pixel at 7/8 brightness and crushes AMV-0 pixels, roughly a third of a typical sprite, to pure black.

Verified against Escape from Monster Manor's zombie.cels: PIXC is 0x3f803f01, so PPMP_0 (all 32 PLUT entries have bit 15 clear) is MS_PIN with SF = 8 and no second source — a pure colour × (AMV+1) / 8. Its AMV values span the full 0–7 range across 4552 opaque pixels. The bundled Message.cel fixture uses the same mode (0x3f003f00).

M2Suite does not model the second-source term: it reads the framebuffer, which a still decode does not have. Sprites that use it are doing blend effects, not carrying colour.

The streaming Cinepak output cel needs no such correction. The 3DO's CPakSubscriber.c builds its LRForm CCB with ccb_PIXC = 0x1F811F00, and both halves work out to ×8/8 — unity. Only the second source differs (PPMP_1 adds the framebuffer). Decoded video frames are passed straight through at their own brightness.

Packed cels

Packed (RLE) cels store rows as a chain: each row begins with a length word, then a run of control bytes selecting literal / repeat / transparent spans.

Finding (verified): 1-bit packed cels use a variable-length row preamble — the offset from the row start to the first control byte is not fixed. Guessing it produces plausible-looking garbage. The reliable test is to validate the row chain: walk all rows using a candidate preamble length and accept it only if every row lands exactly on the next row's start and the last row ends at the end of the data. Verified against Street Fighter's title screens and StarBlade's DANGER label.

Headerless cels

Some files are raw pixel data with no CCB . These are M1 VRAM captures and are stored in LRForm layout: the framebuffer is split into two interleaved halves (even and odd scanlines in separate regions), so the image must be de-interleaved before it looks like anything.


ANIM and cel chains

The frame rate is 16.16 fixed point

Finding (verified). form3do.h describes AnimChunk::frameRate as the "number of 1/60s of a sec to display each frame" and types it int32, which reads like a plain integer. It is not — it is 16.16 fixed point, so the count of 60ths is frameRate / 65536 and the rate is 60 / (frameRate / 65536) fps.

Confirmed across 350+ ANIM files on real discs: every value is a clean 16.16 number between 2.0 and 20.0, and the fractional ones are exactly 60/n, which is what an authoring tool produces when it divides 60 by a target frame rate:

Raw 16.16 = 60/n fps
0x00020000 2.00 60/30 30
0x00026666 2.40 60/25 25
0x00035555 3.33 60/18 18
0x0005745d 5.45 60/11 11
0x00089249 8.57 60/7 7

0x00020000 — 30 fps — is by far the most common. Read as a plain integer it is 131072, which yields a frame time of about 36 minutes.

Two different things both behave as animations:

  1. ANIM files — an explicit animation chunk wrapping a sequence of cels, with a frame count and frame rate in the header.
  2. Cel chains — a file that is simply several CCB chunks concatenated. There is no header saying "this is an animation"; the only signal is that a second CCB follows the first. M2Suite detects this by walking the flat chunk chain and stopping as soon as it sees a second CCB , which keeps the check cheap on large files.

Finding (verified): A PLUT may appear after the pixel data it belongs to, not before. A parser that only applies a palette forward produces a black or miscoloured frame. M2Suite back-fills a trailing PLUT onto the preceding frame. Verified against StarBlade's boom.Anim.

Finding (verified): Real files append trailing data after the last frame — padding, or another asset entirely. Treating a bad chunk size as fatal throws away a perfectly good animation. The rule that works: if at least one frame has already been decoded, stop cleanly at the bad chunk; only error if nothing was decoded at all. Verified against Street Fighter's PL05_ENDING.DAT (26 frames recovered) and syukyakuDEMO.dat (71 frames), both of which previously failed outright with anim chunk '...' has bad size.

Finding (consistent): Some cel chains carry no header at all, so a synthetic one is derived from the first cel's CCB.


IMAG screens and LRForm

IMAG is a full-screen M1 image, typically 16-bit 0555.

Finding (verified): Some IMAG files declare a compression mode in the header but store a complete uncompressed buffer anyway. Trusting the header yields a mostly-black image. The fix is to check whether the payload size already equals width × height × 2 and, if so, take it verbatim. Verified against Yu Yu Hakusho's STOP.IMAG, which went from 99.7% black pixels to a full image.

Both linear and LRForm (interleaved half-framebuffer) layouts occur, and the same de-interleave as for headerless cels applies.


RSRC resource bundles

A 3DO resource archive that wraps several assets in one file.

'RSRC'  header
'RTBL'  resource table
        u32 count
        records[count]   32 bytes each: type tag, id, offset, size, name
payload

Finding (verified): Files with an .AIFF extension are sometimes actually RSRC bundles. Road Rash's Rash.AIFF is an RSRC containing 19 separate AIFF sounds, and a strict AIFF loader rejects it with expected 'FORM', got 'RSRC'. M2Suite detects the RSRC magic at the top of Aiff::load and unwraps to the first AIFF resource.


Audio codecs

Encoding, not just decoding

M2Suite writes these formats as well as reading them, for asset replacement and translation work. Two things are worth recording.

"Bitrate" is not a free parameter. Every 3DO codec is a fixed-ratio scheme with no quality knob, so the bitrate falls out of sampleRate × channels × bitsPerSample — 16, 8 or 4 bits. A UI that offers a bitrate slider is lying about what the format can do; M2Suite reports the resulting rate instead.

The delta codecs must be encoded against their own decoder. SDX2, SQS2 and CBD2 each carry per-channel history, so an encoder that disagrees with the decoder by one LSB does not sound slightly wrong — the error compounds and the file audibly drifts. The encoder here searches both the "exact" (even byte, history discarded) and "delta" (odd byte) forms for every sample and keeps whichever lands closer, which also stops error accumulating when a signal moves faster than the delta table can follow.

Measured round-trip error against a 440 Hz tone with deliberate transients:

Codec Ratio RMS error (of 32000 full scale)
SDX2 2:1 41 (0.13%)
CBD2 2:1 33 (0.10%)
ADP4 4:1 1337 (4.2%)

Container rule: AIFF has no compression tag, so anything other than PCM needs AIFC — and an AIFC must carry an FVER chunk before its COMM.

3DO AIFF/AIFC files use several proprietary codecs identified by the compression tag in the COMM chunk.

Tag Codec Ratio Notes
NONE / (absent) Linear PCM 8 or 16-bit 1:1
SDX2 Square-Delta 2:1 2:1 Squared-delta with sign; per-channel state
SQS2 SDX2 variant 2:1
CBD2 Callisto Block Delta 2:1
ADP4 IMA/DVI ADPCM, 4-bit 4:1 Canonical 89-entry step table
sowt Byte-swapped PCM 1:1 Not yet supported

Finding (verified): AIFF sample rates are stored as an 80-bit IEEE 754 extended float, not an integer. It must be decoded properly; truncating the mantissa gives rates that are close but wrong, and the drift is audible over a long file.

Finding (verified): ADP4 is standard IMA ADPCM with the canonical 89-entry step-size table and the standard 16-entry index-adjust table — not a 3DO-specific variant. Verified against Yu Yu Hakusho's SOUND/*.sc.

Identifying an unknown codec statistically

Technique (verified, and reusable): when you have a byte stream and several candidate interpretations, decode it under each and measure the lag-1 autocorrelation of the resulting samples. Real audio is strongly correlated sample-to-sample; a wrong interpretation is close to noise.

This settled the FILM audio question outright:

Interpretation Lag-1 autocorrelation
Signed 8-bit PCM (what FFmpeg's segafilm assumes) 0.013
Unsigned 8-bit PCM 0.014
SDX2 0.923

No listening test needed — 0.92 versus 0.01 is not ambiguous.


DataStreamer movies

3DO's streaming container, used for in-game movies. A flat sequence of chunks, each:

u32 tag
u32 size        including this 8-byte header
u32 time        stream-tick timestamp
u32 channel
payload
Tag Contents
SHDR Stream header: tick rate, buffer sizing, channel map
FILM Video subchunk (see below)
SNDS Audio subchunk
CTRL Control/marker
FILL Padding

FILM subchunks carry a further sub-tag: FHDR (film header), FRME (keyframe), DFRM (difference frame).

Finding (verified): Files exist that begin with CTRL or FILL rather than SHDR. Requiring a stream header at offset 0 rejects them. M2Suite scans forward, counting payload chunks, and accepts the file if the chunk chain is coherent. Verified against several Need For Speed Movies/*.Stream.

Finding (verified): DFRM difference frames must be decoded, not skipped. A decoder that only handles FRME plays keyframes only and looks like a slideshow — Coven's Intro.stream went from 242 usable frames to 6137 once DFRM was handled.

Finding (verified) — A/V sync: frame timestamps are in stream ticks, and the tick rate in SHDR is not always usable. Deriving the effective rate from the audio track's true decoded duration and then showing the newest frame whose timestamp is due keeps long clips in sync, and tolerates variable frame spacing. Assuming a constant frame rate drifts.


Standalone FILM

The same FILM payload without the DataStreamer wrapper — a QuickTime derivative:

'FILM'  size  ...
  'FDSC'   film description: codec fourcc, width, height
  'STAB'   sample table: per-sample offset, size, timestamp, flags
  'FRME'   the samples themselves

Finding (verified) — the audio is SDX2, not PCM. The generic FILM/Sega FILM description (and FFmpeg's segafilm demuxer) treats 8-bit FILM audio as raw signed PCM. On 3DO discs it is SDX2-compressed at 22050 Hz, mono, decoding to 16-bit. Decoding it as PCM produces loud noise. Established by the autocorrelation test above against Yu Yu Hakusho's MOVIE/MV_01.film.

Finding (consistent): DFRM appears here too, alongside FRME, and both must be accepted as sample types.


Cinepak

The video codec inside almost all 3DO FILM payloads (cvid).

Finding (verified): inter-frame (delta) decoding must follow the canonical algorithm exactly — in particular, the codebook update flags and the "skip" run semantics. An approximation produces progressive block corruption that only becomes obvious dozens of frames in. M2Suite's decoder was rewritten against the reference algorithm after exactly that failure.


MPEG-1 in 3DO containers

Some discs carry MPEG-1 elementary streams inside 3DO containers.

Finding (verified): when an MPEG frame spans several chunks, the continuation chunks use the tag FRM] and — unlike the first chunk — carry no timestamp prefix. Stripping four bytes from every chunk uniformly corrupts the stream. Verified against Oldsmobile's 72130.stream.

Finding (verified): M1VC containers nest. Pontiac's ControlCenterUP.m1c wraps its real MPEG-1 640×480 payload six levels deep. Unwrapping needs to recurse with a depth guard rather than peel one layer.


Opera filesystem

The 3DO disc filesystem, present on M1 and M2 discs.

  • 2048-byte logical blocks; the volume header is at block 0 and carries the magic byte 0x01 plus the ZZZZZ synchronisation pattern.
  • Directories are block-chained; each entry has a type tag, block count, and a fixed-length name field.
  • Filenames may be Shift-JIS on Japanese discs and must not be forced through a Latin-1 conversion.

Finding (verified): M2Suite's extractor was checked byte-identical against an established reference extractor across full discs, which is what makes disc extraction trustworthy enough to build everything else on.

.chd images are not supported — convert to .bin/.cue or .iso first.


M2 PowerPC executables

M2 runs PowerPC 602. Executables are ELF.

Finding (consistent): sections are frequently LZSS-compressed and must be inflated before disassembly, or the listing is noise. M2Suite detects and decompresses these automatically.

The disassembler resolves branch targets and applies symbol names from the ELF symbol table, and a second pass lifts the listing into ANSI-C-like pseudocode.


Alone in the Dark

The 3DO ports of Alone in the Dark 1 and 2 keep the DOS engine's data formats, so everything here is little-endian. This is the most complete set of findings in the project.

PAK archives

u32 offsets[N]        offsets[0] == 0; offsets[1] doubles as the table size,
                      so N = offsets[1] / 4  and  entryCount = N - 1

per entry, at offsets[i]:
  u32 skip            size of the extra header, including this field
  u8  extra[skip - 4]
  u32 compressedSize
  u32 uncompressedSize
  u8  compressionType
  u8  compressionFlags
  u16 padding
  u8  payload[]       at entryOffset + 16 + extraLen + padding

Finding (verified) — the compression is PKWARE DCL "implode", not LZSS. This cost real time. compressionType == 1 was initially assumed to be the LZSS used elsewhere in the engine, and that assumption was empirically disproved before the right answer was found: the first control byte 0x0F decodes as "four literals, then a back-reference to offset 789", which is impossible when only four bytes of output exist. The actual codec is PKWARE Data Compression Library implode, decoded with Mark Adler's public-domain blast. In the compression flags, bit 2 (0x04) means a literal Huffman tree is present and bit 1 (0x02) selects the 8 KiB dictionary.

compressionType == 0 is stored; 4 is a deflate-like variant.

Body (3D model) format

Per the fitd reference implementation (hqr.cpp, createBodyFromPtr):

u16 flags
s16 bbox[6]           ZVX1, ZVX2, ZVY1, ZVY2, ZVZ1, ZVZ2
u16 scratchSize
u8  scratch[scratchSize]
u16 numVertices
s16 vertices[numVertices][3]
                      -- if flags & INFO_ANIM (2):
u16 numGroups
u16 groupOrder[numGroups]
u8  groupRecords[numGroups][recSize]
                      recSize = 0x18 if flags & INFO_OPTIMISE (8), else 0x10
u16 numPrimitives
    primitives[numPrimitives]

Flags: INFO_ANIM = 2, INFO_TORTUE = 4, INFO_OPTIMISE = 8.

Finding (verified) — the field at offset 14 is scratchBufferSize, not the vertex count. An early heuristic that searched for a vertex count near the header found a value that "fit" and produced geometry that looked almost right. It is actually the size of a variable-length scratch buffer that sits before the real count, so every model was being read at the wrong offset. There is no shortcut here: the layout must be walked exactly.

Bone groups — vertices are NOT in model space

Finding (verified) — the single most consequential AITD finding so far.

A body with INFO_ANIM stores its vertices in group-local space. Each group (bone) owns a run of vertices, and every one of them is an offset from the position of that group's base vertex — the vertex it hangs off in its parent. A parser that stops after reading the vertex array gets a mesh whose limbs are all bunched around the origin.

To resolve it, walk the groups in stored order, in place:

for each group g, in the order they appear in the file:
    base = vertices[g.baseVertex]          // already in model space
    for k in 0 .. g.vertexCount-1:
        vertices[g.start + k] += base

Two details are load-bearing:

  • Stored order matters. Groups are laid out parents-first, so by the time a child is reached its base vertex has already been moved into model space and the offsets cascade correctly down the chain.
  • It must be in place. Resolving against a pristine copy of the vertex array breaks every group below the first level, because children need the parent's resolved position, not its original one.

start and baseVertex are stored as byte offsets — divide by 6.

Group record layout, following the group-order table (one u16 per group):

Field Size Notes
start u16 first vertex, byte offset (÷6)
numVertices u16 how many vertices the group owns
baseVertices u16 base vertex, byte offset (÷6)
orgGroup s8 parent group index, −1 at the root
numGroup s8
state.type s16 0 = rotate, 1 = translate, 2 = zoom
state.delta[3] s16 × 3 animation deltas, zero in the bind pose
state.rotateDelta[3] + pad s16 × 4 AITD2 only (INFO_OPTIMISE)

So the record is 0x10 bytes on AITD1 and 0x18 on AITD2. Reading the wrong stride swallows the primitive count and the whole primitive list decodes as garbage.

How it was caught: LISTBOD2.PAK entry 12 (AITD1) is Emily Hartwood. It rendered as a jumble of overlapping shards. Comparing against a reference render of the same body made it obvious the geometry was present but misplaced, not corrupt — which pointed at a transform rather than a parse error. Pinned by regression tests (test_aitd_synth.cpp, testGroupHierarchyResolves).

Primitives

u8 type
Type Name Payload after type
0 Line u8 subType, u8 color, u8 even, u16 points[2]
1 Poly u8 n, u8 subType, u8 color, u16 points[n]
2 Point u8 subType, u8 color, u8 even, u16 points[1]
3 Sphere u8 subType, u8 color, u8 even, u16 size, u16 points[1]
4 Disk (not observed in 3DO data)
5 Cylinder (not observed in 3DO data)
6 BigPoint as Point
7 Zixel as Point
8 PolyTexture8 as Poly
9 PolyTexture9 as Poly, then u8 uv[n][2]
10 PolyTexture10 as Poly, then u8 uv[n][2]

Finding (verified) — vertex indices are stored as byte offsets. Divide each u16 by 6 (three s16 per vertex) to get the vertex index.

Finding (verified) — PolyTexture 9 and 10 append a (u,v) byte pair per point; type 8 does not. These UVs sit inline in the primitive stream, so a parser that skips them desynchronises every following primitive. The symptom is distinctive: models grow long thin spikes shooting off into space, because garbage indices reference distant vertices. This is pinned by a regression test (tests/libm2model_tests/test_aitd_synth.cpp).

Finding (verified) — there is no textured geometry in the 3DO builds. A sweep of every PAK in both games found zero primitives of type 8, 9 or 10 in any real model archive:

Archive Game Bodies Primitive types present
LISTBODY.PAK AITD 1 272 0, 1, 2, 3, 6, 7
LISTBOD2.PAK AITD 1 272 0, 1, 2, 3, 6, 7
4LSTBODY.PAK AITD 2 551 / 553 0, 1, 2, 3, 6, 7

Every polygon is flat-shaded from the palette. So "sample the PolyTexture primitives" has no data to act on for these two ports — the UV parsing is implemented and tested so that a build which does use them (the PC releases, or AITD 3) decodes correctly, but on 3DO there is nothing to sample. The reference implementation fitd reaches the same conclusion from the other direction: it renders PolyTexture primitives with processPrim_Poly, i.e. flat, and never samples a texture.

Which archives hold what

Determined by sweeping every .PAK and counting entries that parse as real geometry (≥4 vertices and ≥2 primitives):

Archive Contents
LISTBODY.PAK, LISTBOD2.PAK, *LSTBODY.PAK 3D models
LISTANIM.PAK, LISTANI2.PAK, 5LSTANIM.PAK Model animations
ETAGE00..15.PAK Floors / rooms
MASK00..15.PAK Camera masks (2D clipping polygons)
ANIM00..15.PAK Room animations
ITD_RESS.PAK, JAP_RESS.PAK Engine resources
LISTSAMP.PAK, chsamp.pak Sound samples
0LSTMAT / 1LSTLIFE / 2LSTTRAK / 3LSTHYB Scripts and tables

Trap: almost any blob can be coerced through the body parser and yield one or two "valid" bodies. MASK PAKs in particular produce a handful of false positives. Identifying a model archive therefore requires a majority of sampled entries to carry real geometry, not merely one.

Pre-rendered backdrops — .pics / .bob / .pad

The room backgrounds. Three container extensions, one payload layout:

u16 width          320 on every file seen
u16 height         200
u8  palette[768]   256 RGB triples, 8 bits per channel
u8  pixels[w*h]    one palette index per pixel
  • .bob — a single un-padded page
  • .pad — a single padded page
  • .pics — one or more pages, each padded

Finding (verified) — by exact arithmetic. 4 + 768 + 320×200 = 64772, which is the exact byte size of CAM8003.BOB. That is not a coincidence, and it confirms the palette size and pixel depth in one step. Decoded output was then visually checked against recognisable rooms.

Finding (verified) — pages pad to a 4 KiB boundary, and dimensions are not fixed. This one was originally got wrong in an instructive way. Every AITD2 backdrop first sampled was 320×200, whose payload of 64772 bytes rounds to exactly 65536 — so "pages are padded to 64 KiB" fitted perfectly and was hard-coded. It is wrong:

File Size Payload Stride Pages
camera00.pics (AITD1) 240×200 48772 49152 5
camera01.pics (AITD2) 320×250 80772 81920 15
camera03.pics (AITD2) 320×200 64772 65536 2

The rule is stride = roundUp(4 + 768 + w×h, 4096), and every page in a file shares the first page's dimensions. Under the 64 KiB assumption both non-320×200 files were rejected outright — a reminder that a constant which fits every sample can still be a coincidence, and that the cheapest defence is a sample whose dimensions differ.

Sound catalogues — LISTSAMP.CAT

Finding (verified). AITD2 ships its sound effects as many complete FORM/AIFF files concatenated into one container, each starting on a 2048-byte boundary (one CD sector) with the gaps zero-padded. LISTSAMP.CAT holds 193 sounds; .FRE and .jpn carry the localised sets.

The sector alignment is the reliable signal. Scanning for the FORM tag at any offset also matches the byte pattern occurring inside sample data — three false positives in LISTSAMP.CAT alone — so only sector boundaries are considered. Walking the chain sequentially does not work either: the file has gaps that are not sounds, and a sequential walk stops after 66 of the 193.

A catalogue is told apart from a single sound structurally: it opens with a FORM whose declared size leaves most of the file unaccounted for.

Rooms and floors — ETAGE*.PAK

Finding (verified). AITD's visuals are the pre-rendered backdrops, so a "room" stores no mesh. What it stores is a set of axis-aligned boxes: colliders the player walks on and bumps into, and triggers that fire scripts. Reassembled, they give a floor's true layout and scale.

Every room of a floor lives in the archive's first entry, behind a u32 offset table; the second entry is camera data. (A later engine revision put one room per entry — AITD-roomviewer switches on the entry count. Both 3DO builds use the table form, with exactly two entries per archive.)

entry 0:
  u32 roomOffsets[]     offsets[0] is BOTH the table size and the first
                        room's offset — the rooms begin immediately after
                        the table. A zero or out-of-range slot ends the list.

per room, relative to its offset:
  u16 colliderOffset    @ 0    relative to the ROOM, not the entry
  u16 triggerOffset     @ 2
  s16 position[3]       @ 4    room origin; multiply by 10 for engine units
  u16 cameraCount       @ 10
  u16 cameraIds[n]      @ 12

at colliderOffset (and likewise at triggerOffset):
  u16 count
  count x 16 bytes:
    s16 lowerX, upperX, lowerY, upperY, lowerZ, upperZ
    s16 id      @ 12
    u16 flags   @ 14

Collider flags: 0x02 underground floor, 0x04 link to another room, 0x08 interactive.

Two traps. The box-list offsets are relative to the room, not to the entry — using entry-relative offsets yields zero rooms rather than wrong ones, which at least fails loudly. And the room position is in units ten times coarser than the box coordinates, so a floor assembled without the ×10 collapses into a heap.

Verified against both 3DO games; AITD1's eight floors yield 1–14 rooms each (ETAGE03: 14 rooms, 178 colliders, 43 triggers, 72 cameras) and render as a recognisable Derceto floor plan.

Derived from tigrouind's AITD-roomviewer (RoomLoader.cs).

Palette

The 256-colour AITD palette is a fixed table, transcribed from AITD_PakEdit's AloneFile::palette and vendored as libs/libm2model/src/AitdPalette.cpp. Primitive color fields index it directly.

Coordinate system

Finding (verified): AITD is Y-down — negative Y is up. A standing character has feet near y = 0 and head near y = -1700. Because screen Y also increases downward, model Y maps straight to screen Y and models render upright with no flip. Exporters targeting Y-up tools (OBJ) must negate Y and Z to preserve handedness; M2Suite's OBJ exporter does.

Entry name databases

AITD_PakEdit ships hand-curated JSON databases (AITD1_CD_PAK_DB.json, AITD1_floppy_PAK_DB.json, …) that name the contents of each archive entry. They are the community's accumulated identification work and turn "Model 12" into "Emily Hartwood".

{ "all_PAKs": {
    "CAMERA02.PAK": {
      "25": { "default_compr": 1, "info": "1st Floor Library (Secret Room)", "type": 2 }
    } } }

Keys are archive filenames, then entry index as a string. info is "?" when the entry has not been identified. The type field is a content class:

type Meaning
0 Unclassified / other
1 Text (in-game messages, readable documents)
2 Background image
3 Floor / room layout
4 Camera data
6 Full-screen image sequence
7 3D model (body)

These databases doubled as an independent check on the renderer. After the bone-group fix, a contact sheet of LISTBOD2.PAK was rendered blind and then compared against the database names: entry 12 "Emily", entry 5 "4 Pane Window", entry 0 "Wooden Chest (Closed), Loft", entry 20 "Indian Cover" — every one matched what had been drawn. Cheap, and much stronger evidence than eyeballing a single model.

M2Suite reads any *PAK_DB.json sitting beside the archive or one level up and uses it to label entries. It does not bundle one — the databases belong to their authors, and are game-specific.

Rendering traps

Two bugs here produced the "broken models" symptom, and neither was a parsing problem:

The declared bounding box is not the geometry. It doubles as a collision volume and can be an order of magnitude larger than the mesh. Framing the camera on it renders the model as a speck in the corner. Fit to the actual vertex extents.

Unreferenced vertices drag the centre. The vertex array contains vertices no primitive references — bone roots and animation helpers, typically sitting at the origin. Including them in the extents pushes models low in the frame and clips them off the bottom edge. Fit only to vertices that a primitive actually uses.

The combination that works: compute the centre from the extents of referenced vertices, then scale by the bounding sphere radius rather than per-axis extents. The sphere radius is rotation-invariant, so the model keeps a constant size and never clips as the camera orbits. Both properties are pinned by regression tests.

Painter's algorithm is not enough. Depth-sorting whole faces renders interpenetrating geometry incorrectly — limbs crossing a torso visibly tear. A per-pixel depth buffer, with depth interpolated across each scanline span, fixes it. Sorting is kept as a cheap first pass.

AITD has no lighting model — do not add one. The artists baked shading into their choice of palette index per face, which is why adjacent faces already read as lit. Applying a light term on top double-shades the model; combined with a two-sided winding test it dimmed most of the mesh, and Emily Hartwood came out near-black next to a reference render of the same body. The faithful mode renders the palette colour unmodified. A neutral grey mode keeps the shading, because there the point is to read silhouette and form rather than to reproduce the original.


ReadySoft FMV — the BNDY container

Finding (verified). Both ReadySoft 3DO titles — Dragon's Lair and Space Ace — ship their full-motion video in one container format, keyed by the BNDY ("boundary") magic:

repeat:
  u32 blockSize             covers the whole block from this word
  u32 flags                 0xF0000000 on Space Ace, 0 on Dragon's Lair
  'BNDY'|'ANDY' u16 index u16 mark  u8 pad[16]    24-byte block header
  'FORM' u32 size 'SASQ' <chunks>   a standard EA-IFF-85 FORM (Space Ace)

blockSize chains the blocks; the first block of a movie is tagged BNDY and every continuation is ANDY. Verified: every one of the 29 Dragon's Lair .DAT files and 21 Space Ace .DAT files parses as such a chain.

Space Ace (SASQ) — decoded

Chunks inside the FORM are 4-byte aligned (not 2 — the 2-byte assumption silently loses every chunk after the first odd-sized one):

Chunk Payload
BPAL big-endian RGB555 palette entries, from index 0
BAC8 keyframe: s16 x, s16 y, u16 w, u16 h then w*h raw 8-bit indices
MVB8 identical layout — a raw sub-rect blit, used for scroll-in strips
MOVE s16 x, s16 y — the viewport origin on the canvas
CELL same rect header, then a standard 3DO packed coded 8-bit cel
COPY zero-length marker
FSND / ASND raw signed 8-bit PCM, 22050 Hz mono

The picture model is a palettised canvas that can be larger than the 288×216 viewport: BAC8/MVB8 supply raw pixels, MOVE pans the view, and each frame's CELL composites a packed cel over whatever is already there (transparent runs leave the canvas untouched). Frames are therefore deltas and must be decoded in order.

The CELL payload is the stock 3DO packed-cel encoding, confirmed by arithmetic rather than assumption: a fully transparent row reads 00 00 BF BF BF BF 9F 00, i.e. a row-offset word of 0 then transparent runs of 64+64+64+64+33 = 288 pixels — exactly the frame width — followed by end-of-row, in (0 + 2) * 4 = 8 bytes. Each row is u16 offset (row length in 32-bit words, minus two) then control bytes whose top two bits select 0 end-of-row, 1 literal run, 2 transparent run, 3 repeat, with the low six bits holding count−1.

Timing. Every ASND is exactly 1838 bytes = 1838/22050 s, so the movies run at 11.996 fps — 12 fps. Frame 0 carries an FSND preload that can exceed a second, so M2Suite paces frames off the running audio-byte total rather than a flat 12 fps (a flat rate drifts ~1.2 s over a 48 s movie).

The audio is raw signed 8-bit PCM rather than SDX2 — decided by measurement: as PCM the track has 0.83 lag-1 autocorrelation and never clips; decoded as SDX2 it falls to 0.79 and saturates both rails.

BAC8 keyframes are interlaced. Even and odd scanlines are two temporally distinct fields woven together, so fast-motion keyframes comb. That is verified, not assumed: within a single file the woven-versus- deinterlaced roughness ratio ranges from 1.1 on near-static frames to 11.0 on fast motion. A structural row-order error would be constant; content dependence means real motion. Weaving reproduces what the format stores and what the 3DO fed its CRT. Two alternative row orders were tested and rejected — the stacked-half reading leaves a seam 4–5× rougher than the surrounding image, and the fields are not a spatial offset of one another.

Dragon's Lair — not decoded

Dragon's Lair's blocks carry flags = 0 and no IFF FORM: the payload begins immediately after the 24-byte block header. The SASQ chunk decoder does not apply, so M2Suite maps the block chain and reports it rather than drawing guessed pixels.

ONE.3DO (Dragon's Lair) is not a BNDY container — a small 5 KB file, likely an index or config — and is left as a hex view.

Macintosh PICT

Finding (verified). Paul.pict (3DO Gamepack) is a genuine Macintosh PICT: a 512-byte all-zero header followed by a QuickDraw picture. The version is read from the record after the header — 0x0011 0x02FF marks PICT v2 (extended QuickDraw); Paul.pict is v2, 240×240. M2Suite identifies it and reports the version and bounding rect; where Qt's image plugins can read the PICT it is shown, but the QuickDraw opcode stream (bitmaps, regions, text) is not decoded here.

Headerless / raw audio

Finding (verified). Some game .aif files are raw samples with no FORM wrapper (3DO Gamepack onside/fb_au/003.aif), which made a strict IFF parser throw. The loader now falls back to treating a non-FORM, non-RSRC input as raw signed 8-bit PCM, mono, 22050 Hz so it plays rather than erroring — a best-effort interpretation of a stripped file, not a claim about a stored header. (Measured autocorrelation of 003.aif is weak under both 8- and 16-bit readings, ~0.22 vs ~0.15, so it may be delta-coded; 8-bit is the better default.)

Crystal Dynamics "bigfile" (Gex)

Big-endian throughout, verified against the real 331 MB GXdata/bigfile:

u32 entryCount
u32 reserved[3]        zero in every container seen
entryCount x {
    u32 id             filename hash — no names are stored
    u32 size
    u32 offset         2048-byte (disc sector) aligned
}
payload

Each entry's offset is the previous entry's end rounded up to 2048, which is what confirms the field order.

Finding (verified) — the format is recursive, and nothing is compressed. This was the whole puzzle. Entries are frequently complete bigfiles themselves, using the identical header:

Top-level entries 244
…of which are containers 158
Nested entries inside them 7,991
Leaves after recursing 8,077 (97.4% of the file)
Nesting depth 2

Flat extraction stops one level short and leaves most of the game's assets sealed inside blobs that look like compressed data because they are opaque. They are not. Every leaf was measured for byte entropy and byte-to-byte correlation and not one is compressed — no entry has the high-entropy, low-correlation signature that compression leaves.

So there is no decompressor to write. The answer to "how do I decompress the bigfile" is that you don't; you recurse.

Detection must be strict. The container check requires the three reserved words to be zero, every offset to be sector-aligned and inside the blob, and the records to account for all but the trailing pad sector. Without the zero-reserved-words guard the parser "finds" containers inside ordinary sample data, and a false positive shreds a real asset into noise — worse than leaving it whole.

54% of the disc is one blob, fifteen times over

Finding (verified by byte comparison.) Fifteen top-level entries are each exactly 12,804,096 bytes, at fifteen distinct offsets, with fifteen distinct name hashes — and all fifteen are byte-identical. That is 171 MB of a 316 MB archive holding one repeated blob.

This is not a bug in the reader: the table has 244 distinct offsets and no two entries share one. It is the classic CD-ROM seek-time optimisation — place a copy of the shared asset bank near each level's data so a 2x drive never has to seek across the disc. Worth knowing before anyone concludes their extraction is duplicating output.

Schema derived from the alpha build — names for the hashes

The release bigfile stores only name hashes, which is why extracted entries are numbered. The Gex alpha (3DO Gamepack GEX_alpha) solves this: instead of one monolithic hash-keyed archive, it keeps a small per- level .big index for each named level directory (kungfu1.dir, jungle2.dir, grave1.dir, …), and those indexes store the plaintext asset names next to their hashes:

u32 count
u32 reserved[3]
count x { u32 hash; u32 field1; u32 field2 }
<name table>   NUL-separated paths, e.g. "kungfu1.dir/map.map",
                                          "paralli/carton1b.par"

Finding (verified). The hashes in the alpha .big files match entries in the release bigfile's nested tree — e.g. 2A39723E is paralli/carton1b.par, and it is a real leaf in the release archive. Cross- referencing the alpha's hash→name map against the release bigfile at every depth names 41 release entries with meaningful paths: level maps (grave1.dir/map.map, cartoon5.dir/map.map), parallax backgrounds (paralli/jungle2b.par), and sound banks. This is the same win as the AITD name database, but derived rather than transcribed.

Consistent (partial). The name hash itself is a 4-byte XOR fold — adjacent level names differ in exactly the hash bytes their changed characters XOR (glue1→glue3 flips only the affected bytes). A brute force over fold parameters reproduces 34 of 77 alpha hashes, so the family is confirmed but the exact permutation is not yet pinned; the direct hash→name table from the alpha is what does the naming today. Fully reversing the fold would name all ~8,000 release leaves, not just the 41 the alpha references.

Sound bank

Each copy of that blob is a container of 516 entries whose ids are a counter (1…0x204) rather than hashes — 442 of them the same size. They are sounds:

s32 unknown       varies, always negative (-1 .. -35248)
u32 length        payload bytes
s32 loopStart     -1 when the sound does not loop
s32 loopEnd       -1
u8  pcm[length]   signed 8-bit

(length + 16) rounded up to 4 equals the entry size for 468 of 516 entries — the remainder are the bank's non-sound assets, and that identity is what tells the two apart. 492 of 516 carry -1/-1 loop points. At 22050 Hz the common 23,499-byte payload is ~1.07 s, and the amplitude envelope varies the way speech does rather than holding steady like a test tone.

The sample rate is not stored, and it is not uniform. 22050 Hz is a default chosen because it makes typical entries about a second long, but a subset of sounds audibly play too fast at that rate — they were authored lower. Nothing in the per-sound header encodes the rate: field[0] looked like a candidate but its high 16 bits are always 0xFFFF and its value tracks the content (identical sounds share it), not a rate. M2Suite therefore exposes the playback rate as a control rather than guessing.

The whole tree is 59% redundant. The 8,077 extracted leaves hold only 729 distinct contents. This is the fifteen-fold level-bank duplication seen from the other side: 441 distinct sounds each appear 15 times (once per level bank), and one appears 750 times. So "the same sound under different names in different folders" is exactly that — byte-identical copies, a CD seek-time optimisation, not variants. A recurring field[0] value is just a recurring sound.

M2Suite's extractor therefore de-duplicates by content hash and sorts leaves into type folders (audio/, images/, video/, geometry/, data/) rather than mirroring the nameless container nesting. On Gex that turns 8,077 files (308 MB) into 729 (126 MB): 415 audio, ~312 data/geometry, 2 images, plus the video kept whole. The on-disc layout was optimised for a 2x CD drive; the extracted layout is optimised for a person browsing it.

Recognition, measured through M2Suite's own sniffer over the recursively extracted tree: 7,020 of 8,077 leaves (86.9%) are sounds, 124 minutes of audio. As a falsification check the decoded output was measured for sample-to-sample correlation — 99.4% score above 0.5, mean 0.831 — so the detector is not simply claiming everything.

Duck TrueMotion video — GXdata/duckart

Each movie is four files sharing a stem:

File Contents
.DUK video data, frames laid end to end
.FRM frame index — big-endian u32 offsets into the .DUK
.HDR 48-byte movie header
.TBL 4096 bytes = 256 entries of 16 (codebook)

Verified: the frame index is strictly ascending and covers the video data in all four movies (614, 917, 423 and 209 frames). Each frame begins with a 12-byte preamble whose bytes 8–11 are the constant magic F7 7F 05 E4, which is what makes a .DUK identifiable on sight.

The .HDR is three big-endian u32 fields, a u16 pair, then eight signed (x,y) motion-vector pairs — the same eight across all four movies: (0,-2) (2,-6) (6,-12) (12,-12) (0,-1) (1,-2) (3,-4) (5,-4).

Consistent (hypothesis) — the sidecar files externalise the codec's tables. This is the promising lead, and it is why FFmpeg's truemotion1 cannot be pointed at these frames directly. Stock TrueMotion 1 selects, per frame, one of a set of built-in tables via two header indices: deltaset (a predictor/delta set) and vectable (a vector codebook). The Gex sidecars look like exactly those tables lifted out of the binary:

  • .HDR ends in eight signed (x,y) pairs — the same eight across all four movies — where TrueMotion 1 keeps its predictor/delta vectors.
  • .TBL is 256 × 16 bytes, the shape of a per-code vector codebook.

This is graded consistent, not verified: the .TBL layout does not yet match TrueMotion 1's variable-length sel_vector_table byte-for-byte (FFmpeg builds that table as, per code, a length byte then len/2 delta pairs), so either the .TBL is a pre-expanded form or the mapping is not one-to-one. Confirming it is the next step.

Verified — the frame body is compressed, not raw. Frame 0's body was rendered as 8-bit grayscale and 16-bit 0555 RGB at every plausible width (64…320); every result is noise. So the data is genuinely vector-quantised — there is no width at which it is a raw bitmap, and the codec must be implemented. Frame data measures ~6.9 bits of entropy per byte, consistent with VQ.

Verified — each frame carries two data regions. The 16-byte preamble is two big-endian length words, the magic, then a u32 (zero on frame 0):

u32 region0_len     e.g. 1518 on BASK frame 0
u32 region1_len     e.g. 6644
u8  magic[4]        F7 7F 05 E4
u32 (zero on the first frame)

Decoded (verified). The two regions are audio then video, and the record layout is:

u32 audio_size
u32 video_size
u8  audio[audio_size]
u8  video[video_size]
u8  sector_padding[]

The decoder is in libm2duck, ported from The Ur-Quan Masters 0.8.0 (src/libs/video/dukvid.c and src/libs/sound/decoders/dukaud.c), which carries a working decoder for exactly this 3DO variant. Because that source is GPL-2.0-or-later, libm2duck is GPL-2.0-or-later — compatible with M2Suite's GPL-3.0, and recorded in THIRD_PARTY_LICENSES.md.

Header (.HDR, 48 bytes).

u32 version              3 selects the version-three decode loop
u32 screen_x_offset      metadata for the game's own compositor
u32 screen_y_offset
u16 width_in_4px_blocks
u16 height_in_4px_blocks
s16 luma_prototypes[8]
s16 chroma_prototypes[8]

Codebook (.TBL, 4096 bytes). 256 vectors of 16 bytes. Byte zero of each vector is the count of meaningful following items; the rest select pairs from the .HDR prototypes. At load time these expand into two signed 256 x 16 delta tables (luma and chroma). The low bit of the last delta is an end-of-sequence marker telling the frame loop to fetch another vector byte. This table is precisely what stock TrueMotion 1 keeps compiled in, which is why FFmpeg and ScummVM cannot decode these payloads directly.

Video subframe. A 16-byte 3DO prefix, then vector-coded picture data from offset 0x10. The first big-endian 16-bit word selects the loop: 0x0300 picks the version-three path, anything else the original. Decoding works in four-pixel-wide blocks, reconstructing vertically paired RGB555 pixels — each 32-bit accumulator holds the upper scanline's pixel in bits 31–16 and the lower scanline's in bits 15–0.

Audio subframe.

u16 magic       0xF77F
u16 numsamples  encoded bytes; two decoded nibbles per byte
u16 tag
u16 index_left
u16 index_right
u8  adpcm[numsamples]

22050 Hz stereo, signed 16-bit after decoding. Nibbles alternate left/right (high nibble first). Predictor state persists across frames while the step indices are re-seeded by every subframe, and the step table differs from canonical IMA ADPCM — the 3DO/UQM values are required.

Timing. A constant 14.622 fps; audio is the sync authority.

Movie Version Size Frames Duration
BASK 1 280x152 614 41.99 s
GEXIALL 1 280x176 917 62.71 s
GEXOALL 1 280x176 423 28.93 s
LOGOTHIN 3 280x184 209 14.29 s

Verified against all four movies: BASK decodes a clean basketball scene, LOGOTHIN (the version-three path) the Crystal Dynamics logo, and every decoded audio track's duration matches the table above.

Reference sources:

  • The Ur-Quan Masters 0.8.0 — dukvid.c / dukaud.c, the working decoder this port is derived from.
  • FFmpeg libavcodec/truemotion1.c — the canonical open TrueMotion 1 decoder; truemotion1_decode_header and gen_vector_table16 show how deltaset/vectable drive the built-in tables that Gex externalises.
  • The Duck Corporation's own TrueMotion "S" SDK, including the tm1.0/ decoder source and a Sega Saturn port, both in the DuckTruemotionS reference tree.

The non-audio leaves (sprites / stages)

Of the 729 distinct leaf contents, 415 are sounds and 314 are not. The non-sound blobs are the level-geometry and sprite candidates, and they carry no magic. Characterised by content:

Kind Distinct Notes
Structured data 288 The bulk. Runs of 0xFF, then mixed index bytes and big-endian floats (46 C0 00 00 = 24576.0) — level geometry and object data.
Sparse tables / tilemaps 14 Mostly zero, small
High-entropy (packed pixels?) 7 Larger, ~7.5+ bits/byte
3DO IMAG images 2 Already decoded elsewhere
Offset table / index 1 The very first leaf

These are individually reachable, listed, and shown with a hex header in the app, but their internal formats are not decoded — genuine reverse-engineering targets rather than transcription.

Other Gex content

Path Contents
LaunchMe, gxduckPlay, gxextra, gxsupprt 3DO M1 ARM executables
System/ 42 more ARM binaries plus 65 3INS DSP instruments
layout.json written by the disc extractor, not game data

The ARM executables all begin E1 A0 00 00 — MOV r0,r0, a no-op, as the entry point. 3DO stores ARM code big-endian, so that byte is byte 0. M2Suite's ARM detector originally tested byte 3 (little-endian ARM habit) and so never fired on a real 3DO executable; opening these files is what exposed it.

Formats that resisted

Documented so nobody repeats the search:

  • Gex leaf assets — the container nesting is solved and every leaf is now reachable and uncompressed, but most leaf payloads are still game-internal formats with no magic. The sound bank above is decoded; the level geometry and sprite data are not.
  • PEBM (Oldsmobile) — the container parses, but it stores no width or height anywhere, so the pixel buffer cannot be laid out. Solving this needs either a companion file that carries the dimensions or a sample whose dimensions are known from elsewhere.

Corrections and additions are welcome — see CONTRIBUTING.md. A finding backed by a reproducible observation is worth more than a plausible theory, and a finding that disproves something on this page is worth most of all.