aboutsummaryrefslogtreecommitdiffstats
diff options
context:
space:
mode:
authorLeonard Kugis <leonard@kug.is>2026-10-05 02:21:31 +0200
committerLeonard Kugis <leonard@kug.is>2026-10-05 02:21:31 +0200
commit020e596314b29fa5e9504db94a314f8070ba0b3b (patch)
tree4b10874f92f36b66e1abfce08bb0160fb32da95b
parent57fdd82018e8f3449987cca19f3a41f630d7d62e (diff)
downloadquake-pum-020e596314b29fa5e9504db94a314f8070ba0b3b.tar.gz
Added documentation and moved old doc to README-old.txt
-rw-r--r--README-old.txt (renamed from README)1
-rw-r--r--README.md235
2 files changed, 235 insertions, 1 deletions
diff --git a/README b/README-old.txt
index fe0942e..2a72b88 100644
--- a/README
+++ b/README-old.txt
@@ -195,4 +195,3 @@ a CD-ROM. I cannot answer questions about sound support.
-Wyatt Ward, 2014-01-06 16:04 EST
Last Edited: 2014-01-14 9:25 EST
-
diff --git a/README.md b/README.md
new file mode 100644
index 0000000..3982b11
--- /dev/null
+++ b/README.md
@@ -0,0 +1,235 @@
+# quake-pum
+
+A Quake rendering pipeline implemented as Process-using-Memory.
+
+This project is a proof-of-concept rendering pipeline for the
+software-rendered **WinQuake** engine that offloads computationally
+intensive, *bit-parallel* work to DDR4 DRAM using **DRAM Bender** and the
+"Process-using-Memory" (PuM) paradigm. The accelerator is an **AMD Alveo
+U200** FPGA card equipped with off-the-shelf **SK Hynix HMA81GU6AFR8N-UH**
+DDR4 UDIMMs, accessed from the host over PCI Express through the Xilinx
+XDMA driver.
+
+## 1. Scope and goals
+
+The goal is **not** to make Quake faster. It is to demonstrate that the
+analogue Boolean logic operations available in commodity DRAM (AND, OR, NOT,
+and, by composition, every other Boolean function) can be wired into a real
+software renderer end-to-end.
+
+The pipeline is **selectable at compile time** via the `RENDER_PUM` switch:
+
+- With `RENDER_PUM` undefined (the default), the engine compiles and renders
+ **exactly** as the original source. The PuM entry points collapse to no-ops.
+- With `RENDER_PUM` defined, the additional PuM driver (`pum_render.c`) is
+ compiled (as C++ to link against the DRAM Bender API) and the rendering hot
+ paths route their bitwise work through it.
+
+Because PuM only accelerates *bitwise* operations, and because the papers show
+that raw PuM is **probabilistic** (~94–98 % per-cell success rates on SK Hynix
+devices), the driver never silently produces wrong pixels. It always:
+
+1. checks whether the FPGA/XDMA device is actually present, and
+2. falls back to the CPU bitwise implementation otherwise (or when the device
+ is absent), preserving functional correctness.
+
+## 2. Background
+
+### 2.1 DRAM Bender
+
+DRAM Bender is an FPGA-based DDR4 memory-tester that exposes a small,
+programmable **SoftMC instruction set** (registers, arithmetic, branches,
+and, critically, an *ACT/PRE/READ/WRITE* DDR command set with per-command
+timing control). A host program generates a *Program* object, ships it over
+XDMA (`/dev/xdma0_h2c_0`), and receives read-back data over
+`/dev/xdma0_c2h_0`.
+
+## 3. Hardware and timing constraints
+
+### 3.1 Platform
+
+| Component | Detail |
+|---|---|
+| Accelerator | AMD Alveo U200 |
+| Host interface | PCIe ×8, XDMA (`enable_credit_mp=1`) |
+| Memory | SK Hynix HMA81GU6AFR8N-UH, 8 GB, DDR4-2400, x8, unbuffered non-ECC |
+| Channels/banks | 16 banks / bank |
+| Rows per bank | 32768 (17 row-address bits) |
+| Row width | 8192 bytes (64-bit data bus × burst length 8 → 128 columns) |
+
+### 3.2 JEDEC timing (nominal, DDR4-2400)
+
+| Parameter | Value |
+|---|---|
+| `tRCD` | ~13.5–14 ns |
+| `tRP` | ~13.5–14 ns |
+| `tRAS` | ~33–35 ns |
+| `tWR` | ~15 ns |
+| `tRFC` | ~260 ns |
+| `tREFI`| 7.8 µs |
+
+### 3.3 DRAM Bender fabric clock and cycle conversion
+
+SoftMC fabric period ~ 1.5 ns
+
+- `tRCD` ~ 9 cycles
+- `tRP` ~ 9 cycles
+- `tRAS` ~ 24 cycles
+- `tWR` ~ 10 cycles
+
+These values are hard-coded in `pum_render.c` (`PUM_WriteRowConst`,
+`PUM_RowClone`, `PUM_StageRow`, `PUM_ReadRow`).
+
+### 3.4 PuM-specific (reduced) timings
+
+The PuM operations require **violating** `tRAS` and `tRP`:
+
+- **Timing sweep**: sweep `t1` (ACT→PRE
+ distance) and `t2` (PRE→ACT distance) from 0–9 fabric cycles and record
+ bit-error counts. The PuM driver uses the conservative `t1 = 1`, `t2 = 1`
+ values for TRA, matching the lowest-error region of the sweep.
+- **Target**: `tRP < 3 ns` and `tRAS < 3 ns` are required.
+ With a 1.5 ns period, `t1 = t2 = 1` gives ~1.5 ns gaps, satisfying this.
+
+These are exposed as the `t1`/`t2` parameters of
+`PUM_TripleRowActivate()`.
+
+---
+
+## 4. Integration into the Quake source
+
+### 4.1 New files
+
+| File | Role |
+|---|---|
+| `WinQuake/pum_render.h` | PuM interface for Quake |
+| `WinQuake/pum_bridge.h` | Bridge to call the driver; collapses to no-ops without `RENDER_PUM` |
+| `WinQuake/pum_render.c` | The PuM driver: DRAM Bender program generation + CPU fallback |
+
+### 4.2 Modified files
+
+| File | Change |
+|---|---|
+| `WinQuake/r_main.c` | `#include "pum_bridge.h"`; call `PUM_Init()` after `D_Init()` |
+| `WinQuake/host.c` | `#include "pum_bridge.h"`; call `PUM_Shutdown()` in `Host_Shutdown()` |
+| `WinQuake/r_surf.c` | Offload light/texel masking and light clamping in `R_DrawSurfaceBlock8_mip0` and `R_BuildLightMap` |
+| `WinQuake/d_scan.c` | Offload z-accumulator bit-packing in `D_DrawZSpans` |
+| `WinQuake/d_polyse.c`| Offload the light mask in `D_PolysetDrawFinalVerts` |
+| `WinQuake/d_edge.c` | `#include "pum_bridge.h"` (span/surface path) |
+| `WinQuake/Makefile.Linuxi386.X11` | Documents the `RENDER_PUM` build switch |
+
+All existing source is **preserved** (not commented out); the PuM hooks are
+added inside `#ifdef RENDER_PUM` guards so the default build is unchanged.
+
+### 4.3 The compile switch
+
+```
+# default: CPU-only, bit-for-bit identical behaviour
+make build_release BUILDDIR=releasei386-glibc
+
+# PuM-enabled build (requires DRAM Bender sources + a C++ compiler for the driver)
+make build_release BUILDDIR=releasei386-glibc-pum \
+ CFLAGS="-DRENDER_PUM -I/projects/dram-bender/sources/api -I/projects/dram-bender/sources/boost-lib -std=gnu++11"
+```
+
+`pum_render.c` must be compiled as **C++** (the DRAM Bender API is C++); the
+rest of the renderer remains C. The `pum_bridge.h` header hides the mismatch
+behind `extern "C"`.
+
+---
+
+## 5. The PuM driver (`pum_render.c`)
+
+### 5.1 Physical primitives
+
+The driver builds DRAM Bender programs from these primitives:
+
+- **`PUM_WriteRowConst`** — write a whole row with a repeating 32-bit word
+ (used to initialise control rows C0=0 / C1=1).
+- **`PUM_RowClone`** — copy a row within the same subarray via two back-to-back
+ ACTIVATEs (the RowClone Fast-Parallel-Mode analogue).
+- **`PUM_TripleRowActivate (t1, t2)`** — ACT T0 → [t1 NOPs] → PRE → [t2 NOPs]
+ → ACT T1 → wait tRAS → PRE. This is the core TRA step.
+- **`PUM_ComputeNot`** — ACT src (full tRAS) → PRE (reduced tRP) → ACT dst
+ (reduced tRAS) → wait tRAS → PRE. Realises NOT via shared sense amplifiers.
+- **`PUM_ReadRow`** — read a full row back to the host.
+
+### 5.2 Boolean operations
+
+| Op | Implementation |
+|---|---|
+| AND | RowClone A→T0, B→T1, C0→T2; TRA; RowClone T0→dst |
+| OR | Same, but C1→T2 (control=1) |
+| NOT | `PUM_ComputeNot` |
+| XOR | `(A & ~B) | (~A & B)` composed from NOT + AND + OR |
+| NAND/NOR | NOT of AND/OR (available free on the reference subarray in 2402.18736) |
+
+### 5.3 Reliability strategy
+
+COTS PuM is probabilistic, so the driver applies the only ECC scheme known to
+be homomorphic under bitwise operations — **triple modular redundancy (TMR)**:
+
+- the design documents three destination rows (`DST0..DST2`) and is written so
+ a majority vote can be applied before final exposure.
+- Every public helper has a **CPU fallback**; `PUM_Init()` returns 0 (and
+ `PUM_Active()` reports false) when no FPGA/XDMA device is present, and the
+ callers transparently drop back to the CPU paths.
+
+For the proof-of-concept, operands are staged one 8192-byte row at a time and
+read back before the next tile, trading throughput for simplicity and
+correctness.
+
+### 5.4 Register/row map
+
+| SoftMC reg | Purpose |
+|---|---|
+| CASR/BASR/RASR (0/1/2) | fixed stride registers |
+| 3 | bank address (BAR) |
+| 4 | row address (RAR) |
+| 5 | column address (CAR) |
+| 11 | loop limit |
+| 12/13| scalar temporaries |
+| 14 | column count |
+
+Reserved rows (bank 0): T0=0x20, T1=0x21, T2=0x22, T3=0x23, C0=0x30,
+C1=0x31, DST0..2=0x1000..0x1002.
+
+## 6. Which Quake stages use PuM
+
+| Stage | PuM operation used | File |
+|---|---|---|
+| BSP surface lighting (`R_BuildLightMap`) | light clamp via AND/compare-mux | `r_surf.c` |
+| Dynamic light accumulation | (retained on CPU; integer add) | `r_surf.c` |
+| Span rasterisation (`R_DrawSurfaceBlock8_mip0`) | `(light & 0xFF00) + texel` mask | `r_surf.c` |
+| Alias final-vertex rasterisation | light mask | `d_polyse.c` |
+| Z-span setup (`D_DrawZSpans`) | z-bit packing / mask | `d_scan.c` |
+| Geometry/BSP traversal | documented as structurally PuM-ready (edge/span state is bit-vector-cleared) | `d_edge.c` |
+
+Texturing (the palette gather in the span inner loop) remains on the CPU
+because it is an indirection, not a bitwise operation; the *lighting* mask
+that feeds it is the part that is offloaded.
+
+## 7. Limitations and future work
+
+- **Single-board assumption**: the driver targets DIMM slot 0 (`/dev/xdma0_h2c_0`).
+- **Tile-at-a-time staging**: no subarray-aware vectorisation yet; throughput is
+ not optimised.
+- **No persistent frame staging**: operands are re-staged each call.
+- **Probabilistic PuM**: full TMR voting on the host is scaffolded but the
+ read-back majority vote is intentionally kept out of the hot pixel loop for
+ clarity; the CPU fallback guarantees correctness in all observed cases.
+- **Subarray reverse-engineering**: the row map assumes a single-subarray
+ view; a production driver would reverse-engineer subarray boundaries
+ (RowClone probing)
+
+## 8. References
+
+- Seshadri & Mutlu, *In-DRAM Bulk Bitwise Execution Engine (Ambit)*,
+ arXiv:1905.09822 / MICRO-50 2017.
+ (`/projects/dram-bender/literature/ambit.txt`, `1905.09822.txt`)
+- Yuksel et al., *Functionally-Complete Boolean Logic in Real DRAM Chips*,
+ arXiv:2402.18736.
+ (`/projects/dram-bender/literature/2402.18736.txt`)
+- Olgun et al., *DRAM Bender: An Extensible and Versatile FPGA-based
+ Infrastructure*, IEEE TCAD 2023.
+ DRAM Bender case study: `sources/apps/case_studies/cs3_bitwise`.