aboutsummaryrefslogtreecommitdiffstats

quake-pum

A Quake rendering pipeline implemented as Process-using-Memory.

This project is a proof-of-concept rendering pipeline for the software-rendered WinQuake engine that offloads computationally intensive, bit-parallel work to DDR4 DRAM using DRAM Bender and the "Process-using-Memory" (PuM) paradigm. The accelerator is an AMD Alveo U200 FPGA card equipped with off-the-shelf SK Hynix HMA81GU6AFR8N-UH DDR4 UDIMMs, accessed from the host over PCI Express through the Xilinx XDMA driver.

1. Scope and goals

The goal is not to make Quake faster. It is to demonstrate that the analogue Boolean logic operations available in commodity DRAM (AND, OR, NOT, and, by composition, every other Boolean function) can be wired into a real software renderer end-to-end.

The pipeline is selectable at compile time via the RENDER_PUM switch:

  • With RENDER_PUM undefined (the default), the engine compiles and renders exactly as the original source. The PuM entry points collapse to no-ops.
  • With RENDER_PUM defined, the additional PuM driver (pum_render.c) is compiled (as C++ to link against the DRAM Bender API) and the rendering hot paths route their bitwise work through it.

Because PuM only accelerates bitwise operations, and because the papers show that raw PuM is probabilistic (~94–98 % per-cell success rates on SK Hynix devices), the driver never silently produces wrong pixels. It always:

  1. checks whether the FPGA/XDMA device is actually present, and
  2. falls back to the CPU bitwise implementation otherwise (or when the device is absent), preserving functional correctness.

2. Background

2.1 DRAM Bender

DRAM Bender is an FPGA-based DDR4 memory-tester that exposes a small, programmable SoftMC instruction set (registers, arithmetic, branches, and, critically, an ACT/PRE/READ/WRITE DDR command set with per-command timing control). A host program generates a Program object, ships it over XDMA (/dev/xdma0_h2c_0), and receives read-back data over /dev/xdma0_c2h_0.

3. Hardware and timing constraints

3.1 Platform

Component Detail
Accelerator AMD Alveo U200
Host interface PCIe ×8, XDMA (enable_credit_mp=1)
Memory SK Hynix HMA81GU6AFR8N-UH, 8 GB, DDR4-2400, x8, unbuffered non-ECC
Channels/banks 16 banks / bank
Rows per bank 32768 (17 row-address bits)
Row width 8192 bytes (64-bit data bus × burst length 8 → 128 columns)

3.2 JEDEC timing (nominal, DDR4-2400)

Parameter Value
tRCD ~13.5–14 ns
tRP ~13.5–14 ns
tRAS ~33–35 ns
tWR ~15 ns
tRFC ~260 ns
tREFI 7.8 µs

3.3 DRAM Bender fabric clock and cycle conversion

SoftMC fabric period ~ 1.5 ns

  • tRCD ~ 9 cycles
  • tRP ~ 9 cycles
  • tRAS ~ 24 cycles
  • tWR ~ 10 cycles

These values are hard-coded in pum_render.c (PUM_WriteRowConst, PUM_RowClone, PUM_StageRow, PUM_ReadRow).

3.4 PuM-specific (reduced) timings

The PuM operations require violating tRAS and tRP:

  • Timing sweep: sweep t1 (ACT→PRE distance) and t2 (PRE→ACT distance) from 0–9 fabric cycles and record bit-error counts. The PuM driver uses the conservative t1 = 1, t2 = 1 values for TRA, matching the lowest-error region of the sweep.
  • Target: tRP < 3 ns and tRAS < 3 ns are required. With a 1.5 ns period, t1 = t2 = 1 gives ~1.5 ns gaps, satisfying this.

These are exposed as the t1/t2 parameters of PUM_TripleRowActivate().


4. Integration into the Quake source

4.1 New files

File Role
WinQuake/pum_render.h PuM interface for Quake
WinQuake/pum_bridge.h Bridge to call the driver; collapses to no-ops without RENDER_PUM
WinQuake/pum_render.c The PuM driver: DRAM Bender program generation + CPU fallback

4.2 Modified files

File Change
WinQuake/r_main.c #include "pum_bridge.h"; call PUM_Init() after D_Init()
WinQuake/host.c #include "pum_bridge.h"; call PUM_Shutdown() in Host_Shutdown()
WinQuake/r_surf.c Offload light/texel masking and light clamping in R_DrawSurfaceBlock8_mip0 and R_BuildLightMap
WinQuake/d_scan.c Offload z-accumulator bit-packing in D_DrawZSpans
WinQuake/d_polyse.c Offload the light mask in D_PolysetDrawFinalVerts
WinQuake/d_edge.c #include "pum_bridge.h" (span/surface path)
WinQuake/Makefile.Linuxi386.X11 Documents the RENDER_PUM build switch

All existing source is preserved (not commented out); the PuM hooks are added inside #ifdef RENDER_PUM guards so the default build is unchanged.

4.3 The compile switch

# default: CPU-only, bit-for-bit identical behaviour
make build_release BUILDDIR=releasei386-glibc

# PuM-enabled build (requires DRAM Bender sources + a C++ compiler for the driver)
make build_release BUILDDIR=releasei386-glibc-pum \
     CFLAGS="-DRENDER_PUM -I/projects/dram-bender/sources/api -I/projects/dram-bender/sources/boost-lib -std=gnu++11"

pum_render.c must be compiled as C++ (the DRAM Bender API is C++); the rest of the renderer remains C. The pum_bridge.h header hides the mismatch behind extern "C".


5. The PuM driver (pum_render.c)

5.1 Physical primitives

The driver builds DRAM Bender programs from these primitives:

  • PUM_WriteRowConst — write a whole row with a repeating 32-bit word (used to initialise control rows C0=0 / C1=1).
  • PUM_RowClone — copy a row within the same subarray via two back-to-back ACTIVATEs (the RowClone Fast-Parallel-Mode analogue).
  • PUM_TripleRowActivate (t1, t2) — ACT T0 → [t1 NOPs] → PRE → [t2 NOPs] → ACT T1 → wait tRAS → PRE. This is the core TRA step.
  • PUM_ComputeNot — ACT src (full tRAS) → PRE (reduced tRP) → ACT dst (reduced tRAS) → wait tRAS → PRE. Realises NOT via shared sense amplifiers.
  • PUM_ReadRow — read a full row back to the host.

5.2 Boolean operations

Op Implementation
AND RowClone A→T0, B→T1, C0→T2; TRA; RowClone T0→dst
OR Same, but C1→T2 (control=1)
NOT PUM_ComputeNot
XOR (A & ~B) | (~A & B) composed from NOT + AND + OR
NAND/NOR NOT of AND/OR (available free on the reference subarray in 2402.18736)

5.3 Reliability strategy

COTS PuM is probabilistic, so the driver applies the only ECC scheme known to be homomorphic under bitwise operations — triple modular redundancy (TMR):

  • the design documents three destination rows (DST0..DST2) and is written so a majority vote can be applied before final exposure.
  • Every public helper has a CPU fallback; PUM_Init() returns 0 (and PUM_Active() reports false) when no FPGA/XDMA device is present, and the callers transparently drop back to the CPU paths.

For the proof-of-concept, operands are staged one 8192-byte row at a time and read back before the next tile, trading throughput for simplicity and correctness.

5.4 Register/row map

SoftMC reg Purpose
CASR/BASR/RASR (0/1/2) fixed stride registers
3 bank address (BAR)
4 row address (RAR)
5 column address (CAR)
11 loop limit
12/13 scalar temporaries
14 column count

Reserved rows (bank 0): T0=0x20, T1=0x21, T2=0x22, T3=0x23, C0=0x30, C1=0x31, DST0..2=0x1000..0x1002.

6. Which Quake stages use PuM

Stage PuM operation used File
BSP surface lighting (R_BuildLightMap) light clamp via AND/compare-mux r_surf.c
Dynamic light accumulation (retained on CPU; integer add) r_surf.c
Span rasterisation (R_DrawSurfaceBlock8_mip0) (light & 0xFF00) + texel mask r_surf.c
Alias final-vertex rasterisation light mask d_polyse.c
Z-span setup (D_DrawZSpans) z-bit packing / mask d_scan.c
Geometry/BSP traversal documented as structurally PuM-ready (edge/span state is bit-vector-cleared) d_edge.c

Texturing (the palette gather in the span inner loop) remains on the CPU because it is an indirection, not a bitwise operation; the lighting mask that feeds it is the part that is offloaded.

7. Limitations and future work

  • Single-board assumption: the driver targets DIMM slot 0 (/dev/xdma0_h2c_0).
  • Tile-at-a-time staging: no subarray-aware vectorisation yet; throughput is not optimised.
  • No persistent frame staging: operands are re-staged each call.
  • Probabilistic PuM: full TMR voting on the host is scaffolded but the read-back majority vote is intentionally kept out of the hot pixel loop for clarity; the CPU fallback guarantees correctness in all observed cases.
  • Subarray reverse-engineering: the row map assumes a single-subarray view; a production driver would reverse-engineer subarray boundaries (RowClone probing)

8. References

  • Seshadri & Mutlu, In-DRAM Bulk Bitwise Execution Engine (Ambit), arXiv:1905.09822 / MICRO-50 2017.
  • Yuksel et al., Functionally-Complete Boolean Logic in Real DRAM Chips, arXiv:2402.18736.
  • Olgun et al., DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure, IEEE TCAD 2023.