Skip to content

Emulate PS2 float semantics for EE FPU results and DIV.S/RSQRT.S - #272

Open
drakolordx7 wants to merge 1 commit into
ran-j:mainfrom
drakolordx7:fix/ps2-float-semantics
Open

drakolordx7 wants to merge 1 commit into
ran-j:mainfrom
drakolordx7:fix/ps2-float-semantics

Conversation

@drakolordx7

Copy link
Copy Markdown

The PS2 EE FPU and the VUs have no Inf/NaN. A value with exponent 255 behaves like a huge normal number, an overflowing result saturates to ±max (0x7F7FFFFF), and x/0 gives ±max with the sign of (numerator xor denominator). The generated code followed IEEE, so Inf/NaN leaked into later maths (for example normalising a zero-length vector).

Changes

  • ps2_runtime_macros.h: new ps2_fclamp / ps2_fdiv / ps2_frsqrt / ps2_vclamp helpers. FPU_ADD_S/SUB_S/MUL_S clamp inputs and result; FPU_DIV_S returns ±max on a zero divisor; FPU_SQRT_S takes sqrt(|x|); new FPU_RSQRT_S; PS2_VADD/VSUB/VMUL clamp every lane (bit-based, so it does not depend on NaN handling in _mm_min_ps on the sse2neon path).
  • fpu_translator.cpp: DIV.S no longer emits copysignf(INFINITY, ...). RSQRT.S was emitted as 1/sqrt(fs); the EE computes fs / sqrt(ft).
  • Tests (ps2_float_semantics_tests.cpp, 8 tests): edge values for each macro, and the translated DIV.S/RSQRT.S code with distinct fd/fs/ft registers.

Found while recompiling Killzone (SCUS-97402): RSQRT.S is used in 382 functions there (vector normalisation), and with the wrong operands normals and lighting came out wrong.

Not included, because open PRs cover them: CVT.W.S truncation/saturation (#215) and the VU0 macro VDIV/VSQRT/VRSQRT Q operations (#198). Expect small textual conflicts with those two in ps2_runtime_macros.h and the test registration.

Notes: generated code has to be regenerated and rebuilt. The clamps add a few integer ops per FPU add/sub/mul; this has not been benchmarked separately.

Tests: ps2x_tests run from the repository root: 444/444 (436/436 on main).

Made by drakolord and assisted with Claude Code.

The EE FPU and the VUs have no Inf/NaN. A value with exponent 255 acts
like a huge normal number, an overflowing result saturates to +-max
(0x7F7FFFFF) and x/0 gives +-max with the sign of (fs xor ft). The
generated code followed IEEE instead, so Inf/NaN leaked into later maths
(a normalisation of a zero-length vector, for example).

- ps2_runtime_macros.h: add ps2_fclamp/ps2_fdiv/ps2_frsqrt/ps2_vclamp.
  FPU_ADD_S/SUB_S/MUL_S clamp their inputs and result, FPU_DIV_S returns
  +-max on a zero divisor, FPU_SQRT_S takes sqrt(|x|), and PS2_VADD/VSUB/
  VMUL clamp every lane. FPU_RSQRT_S is new.
- fpu_translator.cpp: DIV.S no longer emits copysignf(INFINITY, ...) and
  wraps the DZ flag update and the division in one block. RSQRT.S was
  emitted as 1/sqrt(fs); the EE computes fs / sqrt(ft).

Found while recompiling Killzone (SCUS-97402): RSQRT.S appears in 382
functions there (vector normalisation), and with the wrong operands the
normals and lighting came out wrong. Verified with the new unit tests
(edge values for each macro, and the translated DIV.S/RSQRT.S code with
distinct fd/fs/ft), and by running Killzone with the change. Generated
code must be regenerated and rebuilt to pick it up.

CVT.W.S and the VU0 macro VDIV/VSQRT/VRSQRT Q-register operations are
left to ran-j#215 and ran-j#198, which cover them.

Made by drakolord and assisted with Claude Code.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant