diff options
| author | Leonard Kugis <leonard@kug.is> | 2026-10-05 02:21:31 +0200 |
|---|---|---|
| committer | Leonard Kugis <leonard@kug.is> | 2026-10-05 02:21:31 +0200 |
| commit | 020e596314b29fa5e9504db94a314f8070ba0b3b (patch) | |
| tree | 4b10874f92f36b66e1abfce08bb0160fb32da95b | |
| parent | 57fdd82018e8f3449987cca19f3a41f630d7d62e (diff) | |
| download | quake-pum-020e596314b29fa5e9504db94a314f8070ba0b3b.tar.gz | |
Added documentation and moved old doc to README-old.txt
| -rw-r--r-- | README-old.txt (renamed from README) | 1 | ||||
| -rw-r--r-- | README.md | 235 |
2 files changed, 235 insertions, 1 deletions
@@ -195,4 +195,3 @@ a CD-ROM. I cannot answer questions about sound support. -Wyatt Ward, 2014-01-06 16:04 EST Last Edited: 2014-01-14 9:25 EST - diff --git a/README.md b/README.md new file mode 100644 index 0000000..3982b11 --- /dev/null +++ b/README.md @@ -0,0 +1,235 @@ +# quake-pum + +A Quake rendering pipeline implemented as Process-using-Memory. + +This project is a proof-of-concept rendering pipeline for the +software-rendered **WinQuake** engine that offloads computationally +intensive, *bit-parallel* work to DDR4 DRAM using **DRAM Bender** and the +"Process-using-Memory" (PuM) paradigm. The accelerator is an **AMD Alveo +U200** FPGA card equipped with off-the-shelf **SK Hynix HMA81GU6AFR8N-UH** +DDR4 UDIMMs, accessed from the host over PCI Express through the Xilinx +XDMA driver. + +## 1. Scope and goals + +The goal is **not** to make Quake faster. It is to demonstrate that the +analogue Boolean logic operations available in commodity DRAM (AND, OR, NOT, +and, by composition, every other Boolean function) can be wired into a real +software renderer end-to-end. + +The pipeline is **selectable at compile time** via the `RENDER_PUM` switch: + +- With `RENDER_PUM` undefined (the default), the engine compiles and renders + **exactly** as the original source. The PuM entry points collapse to no-ops. +- With `RENDER_PUM` defined, the additional PuM driver (`pum_render.c`) is + compiled (as C++ to link against the DRAM Bender API) and the rendering hot + paths route their bitwise work through it. + +Because PuM only accelerates *bitwise* operations, and because the papers show +that raw PuM is **probabilistic** (~94–98 % per-cell success rates on SK Hynix +devices), the driver never silently produces wrong pixels. It always: + +1. checks whether the FPGA/XDMA device is actually present, and +2. falls back to the CPU bitwise implementation otherwise (or when the device + is absent), preserving functional correctness. + +## 2. Background + +### 2.1 DRAM Bender + +DRAM Bender is an FPGA-based DDR4 memory-tester that exposes a small, +programmable **SoftMC instruction set** (registers, arithmetic, branches, +and, critically, an *ACT/PRE/READ/WRITE* DDR command set with per-command +timing control). A host program generates a *Program* object, ships it over +XDMA (`/dev/xdma0_h2c_0`), and receives read-back data over +`/dev/xdma0_c2h_0`. + +## 3. Hardware and timing constraints + +### 3.1 Platform + +| Component | Detail | +|---|---| +| Accelerator | AMD Alveo U200 | +| Host interface | PCIe ×8, XDMA (`enable_credit_mp=1`) | +| Memory | SK Hynix HMA81GU6AFR8N-UH, 8 GB, DDR4-2400, x8, unbuffered non-ECC | +| Channels/banks | 16 banks / bank | +| Rows per bank | 32768 (17 row-address bits) | +| Row width | 8192 bytes (64-bit data bus × burst length 8 → 128 columns) | + +### 3.2 JEDEC timing (nominal, DDR4-2400) + +| Parameter | Value | +|---|---| +| `tRCD` | ~13.5–14 ns | +| `tRP` | ~13.5–14 ns | +| `tRAS` | ~33–35 ns | +| `tWR` | ~15 ns | +| `tRFC` | ~260 ns | +| `tREFI`| 7.8 µs | + +### 3.3 DRAM Bender fabric clock and cycle conversion + +SoftMC fabric period ~ 1.5 ns + +- `tRCD` ~ 9 cycles +- `tRP` ~ 9 cycles +- `tRAS` ~ 24 cycles +- `tWR` ~ 10 cycles + +These values are hard-coded in `pum_render.c` (`PUM_WriteRowConst`, +`PUM_RowClone`, `PUM_StageRow`, `PUM_ReadRow`). + +### 3.4 PuM-specific (reduced) timings + +The PuM operations require **violating** `tRAS` and `tRP`: + +- **Timing sweep**: sweep `t1` (ACT→PRE + distance) and `t2` (PRE→ACT distance) from 0–9 fabric cycles and record + bit-error counts. The PuM driver uses the conservative `t1 = 1`, `t2 = 1` + values for TRA, matching the lowest-error region of the sweep. +- **Target**: `tRP < 3 ns` and `tRAS < 3 ns` are required. + With a 1.5 ns period, `t1 = t2 = 1` gives ~1.5 ns gaps, satisfying this. + +These are exposed as the `t1`/`t2` parameters of +`PUM_TripleRowActivate()`. + +--- + +## 4. Integration into the Quake source + +### 4.1 New files + +| File | Role | +|---|---| +| `WinQuake/pum_render.h` | PuM interface for Quake | +| `WinQuake/pum_bridge.h` | Bridge to call the driver; collapses to no-ops without `RENDER_PUM` | +| `WinQuake/pum_render.c` | The PuM driver: DRAM Bender program generation + CPU fallback | + +### 4.2 Modified files + +| File | Change | +|---|---| +| `WinQuake/r_main.c` | `#include "pum_bridge.h"`; call `PUM_Init()` after `D_Init()` | +| `WinQuake/host.c` | `#include "pum_bridge.h"`; call `PUM_Shutdown()` in `Host_Shutdown()` | +| `WinQuake/r_surf.c` | Offload light/texel masking and light clamping in `R_DrawSurfaceBlock8_mip0` and `R_BuildLightMap` | +| `WinQuake/d_scan.c` | Offload z-accumulator bit-packing in `D_DrawZSpans` | +| `WinQuake/d_polyse.c`| Offload the light mask in `D_PolysetDrawFinalVerts` | +| `WinQuake/d_edge.c` | `#include "pum_bridge.h"` (span/surface path) | +| `WinQuake/Makefile.Linuxi386.X11` | Documents the `RENDER_PUM` build switch | + +All existing source is **preserved** (not commented out); the PuM hooks are +added inside `#ifdef RENDER_PUM` guards so the default build is unchanged. + +### 4.3 The compile switch + +``` +# default: CPU-only, bit-for-bit identical behaviour +make build_release BUILDDIR=releasei386-glibc + +# PuM-enabled build (requires DRAM Bender sources + a C++ compiler for the driver) +make build_release BUILDDIR=releasei386-glibc-pum \ + CFLAGS="-DRENDER_PUM -I/projects/dram-bender/sources/api -I/projects/dram-bender/sources/boost-lib -std=gnu++11" +``` + +`pum_render.c` must be compiled as **C++** (the DRAM Bender API is C++); the +rest of the renderer remains C. The `pum_bridge.h` header hides the mismatch +behind `extern "C"`. + +--- + +## 5. The PuM driver (`pum_render.c`) + +### 5.1 Physical primitives + +The driver builds DRAM Bender programs from these primitives: + +- **`PUM_WriteRowConst`** — write a whole row with a repeating 32-bit word + (used to initialise control rows C0=0 / C1=1). +- **`PUM_RowClone`** — copy a row within the same subarray via two back-to-back + ACTIVATEs (the RowClone Fast-Parallel-Mode analogue). +- **`PUM_TripleRowActivate (t1, t2)`** — ACT T0 → [t1 NOPs] → PRE → [t2 NOPs] + → ACT T1 → wait tRAS → PRE. This is the core TRA step. +- **`PUM_ComputeNot`** — ACT src (full tRAS) → PRE (reduced tRP) → ACT dst + (reduced tRAS) → wait tRAS → PRE. Realises NOT via shared sense amplifiers. +- **`PUM_ReadRow`** — read a full row back to the host. + +### 5.2 Boolean operations + +| Op | Implementation | +|---|---| +| AND | RowClone A→T0, B→T1, C0→T2; TRA; RowClone T0→dst | +| OR | Same, but C1→T2 (control=1) | +| NOT | `PUM_ComputeNot` | +| XOR | `(A & ~B) | (~A & B)` composed from NOT + AND + OR | +| NAND/NOR | NOT of AND/OR (available free on the reference subarray in 2402.18736) | + +### 5.3 Reliability strategy + +COTS PuM is probabilistic, so the driver applies the only ECC scheme known to +be homomorphic under bitwise operations — **triple modular redundancy (TMR)**: + +- the design documents three destination rows (`DST0..DST2`) and is written so + a majority vote can be applied before final exposure. +- Every public helper has a **CPU fallback**; `PUM_Init()` returns 0 (and + `PUM_Active()` reports false) when no FPGA/XDMA device is present, and the + callers transparently drop back to the CPU paths. + +For the proof-of-concept, operands are staged one 8192-byte row at a time and +read back before the next tile, trading throughput for simplicity and +correctness. + +### 5.4 Register/row map + +| SoftMC reg | Purpose | +|---|---| +| CASR/BASR/RASR (0/1/2) | fixed stride registers | +| 3 | bank address (BAR) | +| 4 | row address (RAR) | +| 5 | column address (CAR) | +| 11 | loop limit | +| 12/13| scalar temporaries | +| 14 | column count | + +Reserved rows (bank 0): T0=0x20, T1=0x21, T2=0x22, T3=0x23, C0=0x30, +C1=0x31, DST0..2=0x1000..0x1002. + +## 6. Which Quake stages use PuM + +| Stage | PuM operation used | File | +|---|---|---| +| BSP surface lighting (`R_BuildLightMap`) | light clamp via AND/compare-mux | `r_surf.c` | +| Dynamic light accumulation | (retained on CPU; integer add) | `r_surf.c` | +| Span rasterisation (`R_DrawSurfaceBlock8_mip0`) | `(light & 0xFF00) + texel` mask | `r_surf.c` | +| Alias final-vertex rasterisation | light mask | `d_polyse.c` | +| Z-span setup (`D_DrawZSpans`) | z-bit packing / mask | `d_scan.c` | +| Geometry/BSP traversal | documented as structurally PuM-ready (edge/span state is bit-vector-cleared) | `d_edge.c` | + +Texturing (the palette gather in the span inner loop) remains on the CPU +because it is an indirection, not a bitwise operation; the *lighting* mask +that feeds it is the part that is offloaded. + +## 7. Limitations and future work + +- **Single-board assumption**: the driver targets DIMM slot 0 (`/dev/xdma0_h2c_0`). +- **Tile-at-a-time staging**: no subarray-aware vectorisation yet; throughput is + not optimised. +- **No persistent frame staging**: operands are re-staged each call. +- **Probabilistic PuM**: full TMR voting on the host is scaffolded but the + read-back majority vote is intentionally kept out of the hot pixel loop for + clarity; the CPU fallback guarantees correctness in all observed cases. +- **Subarray reverse-engineering**: the row map assumes a single-subarray + view; a production driver would reverse-engineer subarray boundaries + (RowClone probing) + +## 8. References + +- Seshadri & Mutlu, *In-DRAM Bulk Bitwise Execution Engine (Ambit)*, + arXiv:1905.09822 / MICRO-50 2017. + (`/projects/dram-bender/literature/ambit.txt`, `1905.09822.txt`) +- Yuksel et al., *Functionally-Complete Boolean Logic in Real DRAM Chips*, + arXiv:2402.18736. + (`/projects/dram-bender/literature/2402.18736.txt`) +- Olgun et al., *DRAM Bender: An Extensible and Versatile FPGA-based + Infrastructure*, IEEE TCAD 2023. + DRAM Bender case study: `sources/apps/case_studies/cs3_bitwise`. |
