1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
|
# quake-pum
A Quake rendering pipeline implemented as Process-using-Memory.
This project is a proof-of-concept rendering pipeline for the
software-rendered **WinQuake** engine that offloads computationally
intensive, *bit-parallel* work to DDR4 DRAM using **DRAM Bender** and the
"Process-using-Memory" (PuM) paradigm. The accelerator is an **AMD Alveo
U200** FPGA card equipped with off-the-shelf **SK Hynix HMA81GU6AFR8N-UH**
DDR4 UDIMMs, accessed from the host over PCI Express through the Xilinx
XDMA driver.
## 1. Scope and goals
The goal is **not** to make Quake faster. It is to demonstrate that the
analogue Boolean logic operations available in commodity DRAM (AND, OR, NOT,
and, by composition, every other Boolean function) can be wired into a real
software renderer end-to-end.
The pipeline is **selectable at compile time** via the `RENDER_PUM` switch:
- With `RENDER_PUM` undefined (the default), the engine compiles and renders
**exactly** as the original source. The PuM entry points collapse to no-ops.
- With `RENDER_PUM` defined, the additional PuM driver (`pum_render.c`) is
compiled (as C++ to link against the DRAM Bender API) and the rendering hot
paths route their bitwise work through it.
Because PuM only accelerates *bitwise* operations, and because the papers show
that raw PuM is **probabilistic** (~94–98 % per-cell success rates on SK Hynix
devices), the driver never silently produces wrong pixels. It always:
1. checks whether the FPGA/XDMA device is actually present, and
2. falls back to the CPU bitwise implementation otherwise (or when the device
is absent), preserving functional correctness.
## 2. Background
### 2.1 DRAM Bender
DRAM Bender is an FPGA-based DDR4 memory-tester that exposes a small,
programmable **SoftMC instruction set** (registers, arithmetic, branches,
and, critically, an *ACT/PRE/READ/WRITE* DDR command set with per-command
timing control). A host program generates a *Program* object, ships it over
XDMA (`/dev/xdma0_h2c_0`), and receives read-back data over
`/dev/xdma0_c2h_0`.
## 3. Hardware and timing constraints
### 3.1 Platform
| Component | Detail |
|---|---|
| Accelerator | AMD Alveo U200 |
| Host interface | PCIe ×8, XDMA (`enable_credit_mp=1`) |
| Memory | SK Hynix HMA81GU6AFR8N-UH, 8 GB, DDR4-2400, x8, unbuffered non-ECC |
| Channels/banks | 16 banks / bank |
| Rows per bank | 32768 (17 row-address bits) |
| Row width | 8192 bytes (64-bit data bus × burst length 8 → 128 columns) |
### 3.2 JEDEC timing (nominal, DDR4-2400)
| Parameter | Value |
|---|---|
| `tRCD` | ~13.5–14 ns |
| `tRP` | ~13.5–14 ns |
| `tRAS` | ~33–35 ns |
| `tWR` | ~15 ns |
| `tRFC` | ~260 ns |
| `tREFI`| 7.8 µs |
### 3.3 DRAM Bender fabric clock and cycle conversion
SoftMC fabric period ~ 1.5 ns
- `tRCD` ~ 9 cycles
- `tRP` ~ 9 cycles
- `tRAS` ~ 24 cycles
- `tWR` ~ 10 cycles
These values are hard-coded in `pum_render.c` (`PUM_WriteRowConst`,
`PUM_RowClone`, `PUM_StageRow`, `PUM_ReadRow`).
### 3.4 PuM-specific (reduced) timings
The PuM operations require **violating** `tRAS` and `tRP`:
- **Timing sweep**: sweep `t1` (ACT→PRE
distance) and `t2` (PRE→ACT distance) from 0–9 fabric cycles and record
bit-error counts. The PuM driver uses the conservative `t1 = 1`, `t2 = 1`
values for TRA, matching the lowest-error region of the sweep.
- **Target**: `tRP < 3 ns` and `tRAS < 3 ns` are required.
With a 1.5 ns period, `t1 = t2 = 1` gives ~1.5 ns gaps, satisfying this.
These are exposed as the `t1`/`t2` parameters of
`PUM_TripleRowActivate()`.
---
## 4. Integration into the Quake source
### 4.1 New files
| File | Role |
|---|---|
| `WinQuake/pum_render.h` | PuM interface for Quake |
| `WinQuake/pum_bridge.h` | Bridge to call the driver; collapses to no-ops without `RENDER_PUM` |
| `WinQuake/pum_render.c` | The PuM driver: DRAM Bender program generation + CPU fallback |
### 4.2 Modified files
| File | Change |
|---|---|
| `WinQuake/r_main.c` | `#include "pum_bridge.h"`; call `PUM_Init()` after `D_Init()` |
| `WinQuake/host.c` | `#include "pum_bridge.h"`; call `PUM_Shutdown()` in `Host_Shutdown()` |
| `WinQuake/r_surf.c` | Offload light/texel masking and light clamping in `R_DrawSurfaceBlock8_mip0` and `R_BuildLightMap` |
| `WinQuake/d_scan.c` | Offload z-accumulator bit-packing in `D_DrawZSpans` |
| `WinQuake/d_polyse.c`| Offload the light mask in `D_PolysetDrawFinalVerts` |
| `WinQuake/d_edge.c` | `#include "pum_bridge.h"` (span/surface path) |
| `WinQuake/Makefile.Linuxi386.X11` | Documents the `RENDER_PUM` build switch |
All existing source is **preserved** (not commented out); the PuM hooks are
added inside `#ifdef RENDER_PUM` guards so the default build is unchanged.
### 4.3 The compile switch
```
# default: CPU-only, bit-for-bit identical behaviour
make build_release BUILDDIR=releasei386-glibc
# PuM-enabled build (requires DRAM Bender sources + a C++ compiler for the driver)
make build_release BUILDDIR=releasei386-glibc-pum \
CFLAGS="-DRENDER_PUM -I/projects/dram-bender/sources/api -I/projects/dram-bender/sources/boost-lib -std=gnu++11"
```
`pum_render.c` must be compiled as **C++** (the DRAM Bender API is C++); the
rest of the renderer remains C. The `pum_bridge.h` header hides the mismatch
behind `extern "C"`.
---
## 5. The PuM driver (`pum_render.c`)
### 5.1 Physical primitives
The driver builds DRAM Bender programs from these primitives:
- **`PUM_WriteRowConst`** — write a whole row with a repeating 32-bit word
(used to initialise control rows C0=0 / C1=1).
- **`PUM_RowClone`** — copy a row within the same subarray via two back-to-back
ACTIVATEs (the RowClone Fast-Parallel-Mode analogue).
- **`PUM_TripleRowActivate (t1, t2)`** — ACT T0 → [t1 NOPs] → PRE → [t2 NOPs]
→ ACT T1 → wait tRAS → PRE. This is the core TRA step.
- **`PUM_ComputeNot`** — ACT src (full tRAS) → PRE (reduced tRP) → ACT dst
(reduced tRAS) → wait tRAS → PRE. Realises NOT via shared sense amplifiers.
- **`PUM_ReadRow`** — read a full row back to the host.
### 5.2 Boolean operations
| Op | Implementation |
|---|---|
| AND | RowClone A→T0, B→T1, C0→T2; TRA; RowClone T0→dst |
| OR | Same, but C1→T2 (control=1) |
| NOT | `PUM_ComputeNot` |
| XOR | `(A & ~B) | (~A & B)` composed from NOT + AND + OR |
| NAND/NOR | NOT of AND/OR (available free on the reference subarray in 2402.18736) |
### 5.3 Reliability strategy
COTS PuM is probabilistic, so the driver applies the only ECC scheme known to
be homomorphic under bitwise operations — **triple modular redundancy (TMR)**:
- the design documents three destination rows (`DST0..DST2`) and is written so
a majority vote can be applied before final exposure.
- Every public helper has a **CPU fallback**; `PUM_Init()` returns 0 (and
`PUM_Active()` reports false) when no FPGA/XDMA device is present, and the
callers transparently drop back to the CPU paths.
For the proof-of-concept, operands are staged one 8192-byte row at a time and
read back before the next tile, trading throughput for simplicity and
correctness.
### 5.4 Register/row map
| SoftMC reg | Purpose |
|---|---|
| CASR/BASR/RASR (0/1/2) | fixed stride registers |
| 3 | bank address (BAR) |
| 4 | row address (RAR) |
| 5 | column address (CAR) |
| 11 | loop limit |
| 12/13| scalar temporaries |
| 14 | column count |
Reserved rows (bank 0): T0=0x20, T1=0x21, T2=0x22, T3=0x23, C0=0x30,
C1=0x31, DST0..2=0x1000..0x1002.
## 6. Which Quake stages use PuM
| Stage | PuM operation used | File |
|---|---|---|
| BSP surface lighting (`R_BuildLightMap`) | light clamp via AND/compare-mux | `r_surf.c` |
| Dynamic light accumulation | (retained on CPU; integer add) | `r_surf.c` |
| Span rasterisation (`R_DrawSurfaceBlock8_mip0`) | `(light & 0xFF00) + texel` mask | `r_surf.c` |
| Alias final-vertex rasterisation | light mask | `d_polyse.c` |
| Z-span setup (`D_DrawZSpans`) | z-bit packing / mask | `d_scan.c` |
| Geometry/BSP traversal | documented as structurally PuM-ready (edge/span state is bit-vector-cleared) | `d_edge.c` |
Texturing (the palette gather in the span inner loop) remains on the CPU
because it is an indirection, not a bitwise operation; the *lighting* mask
that feeds it is the part that is offloaded.
## 7. Limitations and future work
- **Single-board assumption**: the driver targets DIMM slot 0 (`/dev/xdma0_h2c_0`).
- **Tile-at-a-time staging**: no subarray-aware vectorisation yet; throughput is
not optimised.
- **No persistent frame staging**: operands are re-staged each call.
- **Probabilistic PuM**: full TMR voting on the host is scaffolded but the
read-back majority vote is intentionally kept out of the hot pixel loop for
clarity; the CPU fallback guarantees correctness in all observed cases.
- **Subarray reverse-engineering**: the row map assumes a single-subarray
view; a production driver would reverse-engineer subarray boundaries
(RowClone probing)
## 8. References
- Seshadri & Mutlu, *In-DRAM Bulk Bitwise Execution Engine (Ambit)*,
arXiv:1905.09822 / MICRO-50 2017.
- Yuksel et al., *Functionally-Complete Boolean Logic in Real DRAM Chips*,
arXiv:2402.18736.
- Olgun et al., *DRAM Bender: An Extensible and Versatile FPGA-based
Infrastructure*, IEEE TCAD 2023.
|