# quake-pum A Quake rendering pipeline implemented as Process-using-Memory. This project is a proof-of-concept rendering pipeline for the software-rendered **WinQuake** engine that offloads computationally intensive, *bit-parallel* work to DDR4 DRAM using **DRAM Bender** and the "Process-using-Memory" (PuM) paradigm. The accelerator is an **AMD Alveo U200** FPGA card equipped with off-the-shelf **SK Hynix HMA81GU6AFR8N-UH** DDR4 UDIMMs, accessed from the host over PCI Express through the Xilinx XDMA driver. ## 1. Scope and goals The goal is **not** to make Quake faster. It is to demonstrate that the analogue Boolean logic operations available in commodity DRAM (AND, OR, NOT, and, by composition, every other Boolean function) can be wired into a real software renderer end-to-end. The pipeline is **selectable at compile time** via the `RENDER_PUM` switch: - With `RENDER_PUM` undefined (the default), the engine compiles and renders **exactly** as the original source. The PuM entry points collapse to no-ops. - With `RENDER_PUM` defined, the additional PuM driver (`pum_render.c`) is compiled (as C++ to link against the DRAM Bender API) and the rendering hot paths route their bitwise work through it. Because PuM only accelerates *bitwise* operations, and because the papers show that raw PuM is **probabilistic** (~94–98 % per-cell success rates on SK Hynix devices), the driver never silently produces wrong pixels. It always: 1. checks whether the FPGA/XDMA device is actually present, and 2. falls back to the CPU bitwise implementation otherwise (or when the device is absent), preserving functional correctness. ## 2. Background ### 2.1 DRAM Bender DRAM Bender is an FPGA-based DDR4 memory-tester that exposes a small, programmable **SoftMC instruction set** (registers, arithmetic, branches, and, critically, an *ACT/PRE/READ/WRITE* DDR command set with per-command timing control). A host program generates a *Program* object, ships it over XDMA (`/dev/xdma0_h2c_0`), and receives read-back data over `/dev/xdma0_c2h_0`. ## 3. Hardware and timing constraints ### 3.1 Platform | Component | Detail | |---|---| | Accelerator | AMD Alveo U200 | | Host interface | PCIe ×8, XDMA (`enable_credit_mp=1`) | | Memory | SK Hynix HMA81GU6AFR8N-UH, 8 GB, DDR4-2400, x8, unbuffered non-ECC | | Channels/banks | 16 banks / bank | | Rows per bank | 32768 (17 row-address bits) | | Row width | 8192 bytes (64-bit data bus × burst length 8 → 128 columns) | ### 3.2 JEDEC timing (nominal, DDR4-2400) | Parameter | Value | |---|---| | `tRCD` | ~13.5–14 ns | | `tRP` | ~13.5–14 ns | | `tRAS` | ~33–35 ns | | `tWR` | ~15 ns | | `tRFC` | ~260 ns | | `tREFI`| 7.8 µs | ### 3.3 DRAM Bender fabric clock and cycle conversion SoftMC fabric period ~ 1.5 ns - `tRCD` ~ 9 cycles - `tRP` ~ 9 cycles - `tRAS` ~ 24 cycles - `tWR` ~ 10 cycles These values are hard-coded in `pum_render.c` (`PUM_WriteRowConst`, `PUM_RowClone`, `PUM_StageRow`, `PUM_ReadRow`). ### 3.4 PuM-specific (reduced) timings The PuM operations require **violating** `tRAS` and `tRP`: - **Timing sweep**: sweep `t1` (ACT→PRE distance) and `t2` (PRE→ACT distance) from 0–9 fabric cycles and record bit-error counts. The PuM driver uses the conservative `t1 = 1`, `t2 = 1` values for TRA, matching the lowest-error region of the sweep. - **Target**: `tRP < 3 ns` and `tRAS < 3 ns` are required. With a 1.5 ns period, `t1 = t2 = 1` gives ~1.5 ns gaps, satisfying this. These are exposed as the `t1`/`t2` parameters of `PUM_TripleRowActivate()`. --- ## 4. Integration into the Quake source ### 4.1 New files | File | Role | |---|---| | `WinQuake/pum_render.h` | PuM interface for Quake | | `WinQuake/pum_bridge.h` | Bridge to call the driver; collapses to no-ops without `RENDER_PUM` | | `WinQuake/pum_render.c` | The PuM driver: DRAM Bender program generation + CPU fallback | ### 4.2 Modified files | File | Change | |---|---| | `WinQuake/r_main.c` | `#include "pum_bridge.h"`; call `PUM_Init()` after `D_Init()` | | `WinQuake/host.c` | `#include "pum_bridge.h"`; call `PUM_Shutdown()` in `Host_Shutdown()` | | `WinQuake/r_surf.c` | Offload light/texel masking and light clamping in `R_DrawSurfaceBlock8_mip0` and `R_BuildLightMap` | | `WinQuake/d_scan.c` | Offload z-accumulator bit-packing in `D_DrawZSpans` | | `WinQuake/d_polyse.c`| Offload the light mask in `D_PolysetDrawFinalVerts` | | `WinQuake/d_edge.c` | `#include "pum_bridge.h"` (span/surface path) | | `WinQuake/Makefile.Linuxi386.X11` | Documents the `RENDER_PUM` build switch | All existing source is **preserved** (not commented out); the PuM hooks are added inside `#ifdef RENDER_PUM` guards so the default build is unchanged. ### 4.3 The compile switch ``` # default: CPU-only, bit-for-bit identical behaviour make build_release BUILDDIR=releasei386-glibc # PuM-enabled build (requires DRAM Bender sources + a C++ compiler for the driver) make build_release BUILDDIR=releasei386-glibc-pum \ CFLAGS="-DRENDER_PUM -I/projects/dram-bender/sources/api -I/projects/dram-bender/sources/boost-lib -std=gnu++11" ``` `pum_render.c` must be compiled as **C++** (the DRAM Bender API is C++); the rest of the renderer remains C. The `pum_bridge.h` header hides the mismatch behind `extern "C"`. --- ## 5. The PuM driver (`pum_render.c`) ### 5.1 Physical primitives The driver builds DRAM Bender programs from these primitives: - **`PUM_WriteRowConst`** — write a whole row with a repeating 32-bit word (used to initialise control rows C0=0 / C1=1). - **`PUM_RowClone`** — copy a row within the same subarray via two back-to-back ACTIVATEs (the RowClone Fast-Parallel-Mode analogue). - **`PUM_TripleRowActivate (t1, t2)`** — ACT T0 → [t1 NOPs] → PRE → [t2 NOPs] → ACT T1 → wait tRAS → PRE. This is the core TRA step. - **`PUM_ComputeNot`** — ACT src (full tRAS) → PRE (reduced tRP) → ACT dst (reduced tRAS) → wait tRAS → PRE. Realises NOT via shared sense amplifiers. - **`PUM_ReadRow`** — read a full row back to the host. ### 5.2 Boolean operations | Op | Implementation | |---|---| | AND | RowClone A→T0, B→T1, C0→T2; TRA; RowClone T0→dst | | OR | Same, but C1→T2 (control=1) | | NOT | `PUM_ComputeNot` | | XOR | `(A & ~B) | (~A & B)` composed from NOT + AND + OR | | NAND/NOR | NOT of AND/OR (available free on the reference subarray in 2402.18736) | ### 5.3 Reliability strategy COTS PuM is probabilistic, so the driver applies the only ECC scheme known to be homomorphic under bitwise operations — **triple modular redundancy (TMR)**: - the design documents three destination rows (`DST0..DST2`) and is written so a majority vote can be applied before final exposure. - Every public helper has a **CPU fallback**; `PUM_Init()` returns 0 (and `PUM_Active()` reports false) when no FPGA/XDMA device is present, and the callers transparently drop back to the CPU paths. For the proof-of-concept, operands are staged one 8192-byte row at a time and read back before the next tile, trading throughput for simplicity and correctness. ### 5.4 Register/row map | SoftMC reg | Purpose | |---|---| | CASR/BASR/RASR (0/1/2) | fixed stride registers | | 3 | bank address (BAR) | | 4 | row address (RAR) | | 5 | column address (CAR) | | 11 | loop limit | | 12/13| scalar temporaries | | 14 | column count | Reserved rows (bank 0): T0=0x20, T1=0x21, T2=0x22, T3=0x23, C0=0x30, C1=0x31, DST0..2=0x1000..0x1002. ## 6. Which Quake stages use PuM | Stage | PuM operation used | File | |---|---|---| | BSP surface lighting (`R_BuildLightMap`) | light clamp via AND/compare-mux | `r_surf.c` | | Dynamic light accumulation | (retained on CPU; integer add) | `r_surf.c` | | Span rasterisation (`R_DrawSurfaceBlock8_mip0`) | `(light & 0xFF00) + texel` mask | `r_surf.c` | | Alias final-vertex rasterisation | light mask | `d_polyse.c` | | Z-span setup (`D_DrawZSpans`) | z-bit packing / mask | `d_scan.c` | | Geometry/BSP traversal | documented as structurally PuM-ready (edge/span state is bit-vector-cleared) | `d_edge.c` | Texturing (the palette gather in the span inner loop) remains on the CPU because it is an indirection, not a bitwise operation; the *lighting* mask that feeds it is the part that is offloaded. ## 7. Limitations and future work - **Single-board assumption**: the driver targets DIMM slot 0 (`/dev/xdma0_h2c_0`). - **Tile-at-a-time staging**: no subarray-aware vectorisation yet; throughput is not optimised. - **No persistent frame staging**: operands are re-staged each call. - **Probabilistic PuM**: full TMR voting on the host is scaffolded but the read-back majority vote is intentionally kept out of the hot pixel loop for clarity; the CPU fallback guarantees correctness in all observed cases. - **Subarray reverse-engineering**: the row map assumes a single-subarray view; a production driver would reverse-engineer subarray boundaries (RowClone probing) ## 8. References - Seshadri & Mutlu, *In-DRAM Bulk Bitwise Execution Engine (Ambit)*, arXiv:1905.09822 / MICRO-50 2017. - Yuksel et al., *Functionally-Complete Boolean Logic in Real DRAM Chips*, arXiv:2402.18736. - Olgun et al., *DRAM Bender: An Extensible and Versatile FPGA-based Infrastructure*, IEEE TCAD 2023.