Over the PCIe link: what the host gets
Every ET-SoC-1 page so far measured the chip from the inside. This one times the way in: the PCIe Gen4 x8 link between the host and the card, per direction by its line rate (16 GT/s on 8 lanes with 128b/130b coding), which had never been timed on these cards. A large copy crosses it at from host to card and back, the same on all three cards to within . A program gets less (), because the runtime first copies its buffer into a bounce buffer with the host's own memcpy. A 4 KB copy issued now and then takes on average from issue to completion, and an empty kernel ; most of that is the runtime polling for the answer, not the card. The numbers were predicted before the runs: .
Terms used on this page
H2D is host to device (the card reads host memory over the link), D2H device to host.
Staged is the path every program takes: memcpyHostToDevice copies the user's buffer into a
DMA-able bounce buffer (CMA memory the driver pins), then the card's PCIe DMA engine moves that buffer; D2H
does the same in reverse. DMA-only runs the same DMA commands with the bounce copy replaced by a no-op,
through the API's own cmaCopyFunction argument: the bytes still cross the link, but the host copy is
gone, which is what pinned memory buys on a GPU (its bytes are not delivered, so it is for timing only). Sizes are
binary (1 MB = 220 B); bandwidth is decimal (1 GB/s = 109 B/s), as the link figure is. A
shire is a group of 32 cores; the chip runs kernels on 32 of them.
1. Bandwidth against transfer size
Each run copies every size from 4 KB to 256 MB, doubling, in both directions and both ways (staged and DMA-only), the four variants in a random order in every repeat; the card's value is the mean of its five runs' medians. The DMA-only line is the link and the card's DMA engine alone. It passes half its large-copy rate at and 90% of it at . Smaller copies use the link less and less, because each one also pays a fixed cost of a tenth of a millisecond or more (section 3).
The DMA-only values at every size, per card
GB/s, the mean of five runs; each value's 99% interval is in the chart's tooltip and in pcie.json.
2. What the bounce copy costs
A program has no pinned-memory call in this runtime: every copy of user memory is staged. For a large copy the
runtime fills the bounce buffer in chunks of up to with a plain memcpy on its thread pool
and enables the DMA command only when all of its chunks are copied (MemcpyH2DAction.cpp in
esperanto-tools-libs; a copy back runs the DMA first, then copies out), so the host copy and the DMA run one after
the other. The staged rate is then close to what the two in series give,
1 / (1/DMA + 1/memcpy), with the host's memcpy measured by the same program on the same host.
That is why the cards stage at different rates over the same link: the hosts' own memcpy of 256 MB runs at
.
Put each card's staged rate against its host's memcpy, and every card lands on the curve of the two copies in series; the drop from the DMA-only tick above it to the point is what the bounce copy costs:
The table, card by card
3. Small copies: the runtime's polling sets the latency
A 4 KB copy spends well under a microsecond on the wire. From issue to completion, after a random idle gap, 98% of them took
(every card pooled), and the spread is not noise. The runtime's response thread polls the card's completion queue; it
sleeps 50 µs between polls while commands are in flight and 500 µs when none are
(ResponseReceiver.cpp in esperanto-tools-libs, the constants
kResponsePollingIntervalWithEventsOnFly and kResponsePollingIntervalNoEventsOnFly).
A copy issued while it is in the long sleep waits for that sleep to end. The chart pools every 4 KB copy of the
five runs: back to back (issued as soon as the previous one returned), or after a random idle gap of 0 to 1 ms,
as a program that copies now and then would.
4. Starting a kernel
The probe's kernel does nothing: each hart enters, returns 0 and goes back to the firmware. One launch, waited for, measures the whole path: the command through the submission queue, the master shire's dispatch to the compute shires, their start and return, the completion back to the host, and the runtime noticing it. A hundred launches queued at once, each with the device-side barrier, measure the card's side alone, because the host's polling then overlaps the work.
Set the two side by side, for a launch and for a small copy issued now and then, and most of the time of one operation is the wait for the runtime's response thread to wake, not the work:
The launch table
5. Several transfers at once
Each stream here moves two 64 MB copies, DMA-only. Without a barrier the runtime sends both commands at once and
the card's DMA worker gives each its own channel (it has four read and four write channels, one command per
channel: pcie_dma.h, dmaw.c in the firmware); with the barrier flag on every copy, one
command of the stream runs at a time. The pilot run found the surprise this section is built around: two H2D
commands in flight move less than one.
Which commands collide (29 September)
A pre-registered follow-up, experiment E55 (the
hub's rungs 34 and 35), asked why. Its rules were set in development on aifoundry1's card 1 (28 September,
23:47–23:55 PDT), frozen (tools/claims-v3/pcie2/PREREG.md, sha256 f632d6b3…), and validated on
aifoundry3 (29 September, 00:05–00:27 PDT, five passes), then run the same way on aifoundry2 as a third card that
evening (16:59–17:21 PDT, five passes, after the DVFS validation there ended). It timed two DMA-only copies per stream of 1, 4, 16 and 64 MB:
both commands in flight, one at a time, and one command in each of two streams. Only two commands of one stream
collide. Two host-to-card commands in flight in one stream moved 0.49 of one at 64 MB on aifoundry3 (a rate ratio of 0.48 from
the fit over 1–64 MB), and 0.49 again with each command split into eight DMA elements; one command in each of two
streams moved 1.01 of one. aifoundry2 gave 0.50 (a rate ratio of 0.49), 0.50 with eight elements and 1.01. A read engine or an IOMMU shared by the two commands would have halved that case too, so both are
refuted, and so are a fixed cost per overlapping command, a loss that sets in only after a long overlap, and a loss per
element. Development on card 1 and the run on aifoundry2 gave the same verdicts. So two streams move as much as one command alone, no more, and two
commands in one stream move half as much. Card to host, two commands in flight moved 1.10 of one at 64 MB on both cards; at 1–16 MB
the passes spread too wide to decide on aifoundry3 (0.78–0.94), and on aifoundry2 two 4 MB commands moved 0.88 of one,
below the 0.95 predicted, a prediction no account of the collision rests on. Why one stream's two commands collide is not established.
Where a host copy lands on the card
The same experiment asked where the bytes of a host-to-card copy are written: into the L3 slice that is each line's
home, or past it to DRAM. A 4 MB buffer was written by the normal (staged) copy; then one hart on shire 0 timed the first
load of 4,096 of its lines, one at a time, in a kernel launched without the L3 flush. Two references were timed the same
way: every line from DRAM (a median of 304 cycles) and every line from the L3 (175 cycles). A host write lands in the
L3. After the copy, 99.5% of the lines read at L3 latency on aifoundry3 and 99.6% on aifoundry2 (99.9% on card 1 in
development), whether or not the L3 held the buffer when the host wrote it, and every value read was the one last written. The rival accounts,
an L3 that updates only lines it already holds, and writes that go to DRAM with the L3's copy invalidated or left stale,
are refuted. A kernel that reads a buffer the host has just copied therefore starts from the L3, not from DRAM; that was
tested for a 4 MB buffer only. Data: docs/reports/data/2026-09-29-pcie2/ (its README and
results.md, every prediction with its verdict).
6. The predictions, and how they fared
The predictions were written down before any timed transfer (docs/reports/data/2026-09-27-pcie/PREREG.md,
its sha256 in the data), from the link's figure, what GPUs achieve on such links, and the runtime's source. A
prediction passes on a card when that card's 99% interval over its five runs lies inside the predicted range, fails
when it lies wholly outside, and is inconclusive otherwise.
7. What this does and does not show
- The runtime API, not the raw link. Every number is what a program gets through
libetrt: its command queues, its bounce buffer and its polling are part of each figure. The DMA-only figures remove only the bounce copy. The link's own protocol efficiency (payload size, the IOMMU, read completions) cannot be separated from the DMA engine's without a PCIe analyser or root access to the configuration space; the maximum payload size was not readable as a normal user. - DMA-only is a stand-in for pinned memory. It moves the same bytes over the link from the same bounce buffer, but a program cannot use it to deliver data; the staged figures are the ones a program gets today.
- One host per card. The cards sit in different hosts (), with different
runtime builds (aifoundry2 the stock build, aifoundry3 a patched one, aifoundry1 a fork:
docs/findings/14-card-behaviour.md), all with the IOMMU in translated mode (hosts.txtin the data). aifoundry3's host copies memory at about half the others' rate, and it is the one host whose memory runs on a single channel (section 2). The DMA-only figures agree across them; the staged ones follow each host's memcpy, and the latency figures follow each build's polling constants. - Unexplained: A 256 MB copy is two 128 MB DMA elements in one command on one channel; why card-to-host dips in the middle sizes was not established.
- The card's clock. The telemetry before and after every run found . Launch latency runs on the master shire's firmware and may change with its clock; the transfer rates were not tested at another clock.
- Five runs per card in one afternoon (); no process held a card longer than , but the card's lock, taken for a whole run of six processes, was held for up to : longer than the lab's rule of 10 s for holding a device, which each process kept. Pilot runs before the schedule (one on each card, two on aifoundry2) are kept in the raw data and left out of every figure; they changed the probe in two ways, described in the method.
Method, and how to reproduce it
The probe. workloads/pciebench is a standalone host program and an empty device kernel in the
et-testdrive style, using only the runtime API. One test per process, each under timeout 10 with an
8.5 s budget and the card's lock held for the whole run: info (device properties, DMA limits, the
negotiated link from sysfs), bw (the sweep, with a 1 MB there-and-back data check on the staged path,
which passed in every run: ), lat (64 B and 4 KB round trips back to back and
after random gaps, 200 pipelined 4 KB copies, an idle-stream wait), launch (150 single launches and three
batches of 100 per shire mask, twice), conc (eight configurations, three trials) and
hostcopy (the host's memcpy, no card). A telemetry sample before and after each run records the die
temperature and clocks.
What the pilots changed. The first pilot on aifoundry2 ran the four copy variants in a fixed order, and
each variant locked onto one of the runtime's two polling modes (one variant always about 0.11 ms, the next always
about 0.58 ms at the same size: the round-trip medians in raw/aifoundry2/pilot1/lat.out). The schedule's
probe shuffles the order in every repeat of the bandwidth sweep, and adds 4 KB copies after random gaps, also
shuffled. Its back-to-back round trips (section 3's table, and predictions P6 and P7) kept the fixed order, host to
card staged, then DMA-only, then card to host staged and DMA-only (workloads/pciebench/host/main.cpp,
testLat), so they can still lock each variant onto one mode. The same pilot showed two queued H2D commands at half the rate of one, so the concurrency test gained
the barrier configurations (the /ser rows). Neither change touched a prediction.
Reduction. workloads/pciebench/reduce_pcie.py: a run's value for a cell is the median of its
repeats, a card's value the mean over its five runs with a 99% t-interval (t = 4.604 for five runs).
cmake -S workloads/pciebench -B build/pciebench -DCMAKE_PREFIX_PATH=/opt/et -Wno-dev && nice cmake --build build/pciebench -j4
TREE=$PWD workloads/pciebench/run_pcie.sh 1 # one run on this host's card (V3_DEVICE=1 on aifoundry1)
workloads/pciebench/schedule.sh 1 5 # from aifoundry2: five rounds on the three cards
python3 workloads/pciebench/reduce_pcie.py docs/reports/data/2026-09-27-pcie/raw \
--prereg docs/reports/data/2026-09-27-pcie/PREREG.md \
--out docs/reports/data/2026-09-27-pcie/pcie.json --md docs/reports/data/2026-09-27-pcie/results.md
python3 scripts/build-report.py pcie-link docs/reports/data/2026-09-27-pcie/pcie.json \
docs/reports/2026-09-27-et-soc1-pcie-link.html
Data: docs/reports/data/2026-09-27-pcie/ (its README lists every file).
Version history and provenance
First written 27 September 2026, from the runs of that afternoon; later that day, the staged path drawn against the hosts' memcpy (section 2) and the time of one launch or small copy split into its queued cost and the wait for the response thread (section 4); 28 September, the hosts' memory channels (section 2), read as a user the night before, where the page had said reading them needs root; 29 September, which commands collide and where a host copy lands (section 5), from E55; the page had left the first open. 30 September: E55's third card, aifoundry2, in section 5. Every number on this page is
computed by its script from docs/reports/data/2026-09-27-pcie/pcie.json, except section 5's two follow-ups, whose figures are quoted from docs/reports/data/2026-09-29-pcie2/results.md (aifoundry3 and aifoundry2) and dev-aifoundry1-c1.md (card 1, development), and the link figure's
arithmetic (16 GT/s × 8 lanes × 128/130 ÷ 8, also in the data) and the runtime's and firmware's constants, which
are cited to their source files above.
8. Related reports
- The ET-SoC-1, interactively — the chip schematic whose host flow these numbers fill in.
- Memory hierarchy — the card's DRAM and caches, for scale against the link.
- Ridge points — the roofline, whose host-link ridge still uses the link's figure, not these measured rates.
- Limits of observability — the hub: every report, and what the meters cannot see.
- docs/findings/ — the findings index and the experiment register.