Over the PCIe link: what the host gets

27 September 2026, with a follow-up on 28–29 September (section 5) · three ET-SoC-1 cards, aifoundry2, aifoundry3 and aifoundry1 card 1, five runs on each, 12 minutes apart · a new probe, workloads/pciebench, through the runtime API only · part of the ET-SoC-1 measurement reports

Every ET-SoC-1 page so far measured the chip from the inside. This one times the way in: the PCIe Gen4 x8 link between the host and the card, per direction by its line rate (16 GT/s on 8 lanes with 128b/130b coding), which had never been timed on these cards. A large copy crosses it at from host to card and back, the same on all three cards to within . A program gets less (), because the runtime first copies its buffer into a bounce buffer with the host's own memcpy. A 4 KB copy issued now and then takes on average from issue to completion, and an empty kernel ; most of that is the runtime polling for the answer, not the card. The numbers were predicted before the runs: .

Terms used on this page

H2D is host to device (the card reads host memory over the link), D2H device to host. Staged is the path every program takes: memcpyHostToDevice copies the user's buffer into a DMA-able bounce buffer (CMA memory the driver pins), then the card's PCIe DMA engine moves that buffer; D2H does the same in reverse. DMA-only runs the same DMA commands with the bounce copy replaced by a no-op, through the API's own cmaCopyFunction argument: the bytes still cross the link, but the host copy is gone, which is what pinned memory buys on a GPU (its bytes are not delivered, so it is for timing only). Sizes are binary (1 MB = 220 B); bandwidth is decimal (1 GB/s = 109 B/s), as the link figure is. A shire is a group of 32 cores; the chip runs kernels on 32 of them.

Host to card, DMA-only, 256 MB
Card to host, DMA-only, 256 MB
Host to card, as a program copies, 256 MB
staged through the bounce buffer; set by the host's memcpy
An empty kernel, launch to completion

1. Bandwidth against transfer size

Each run copies every size from 4 KB to 256 MB, doubling, in both directions and both ways (staged and DMA-only), the four variants in a random order in every repeat; the card's value is the mean of its five runs' medians. The DMA-only line is the link and the card's DMA engine alone. It passes half its large-copy rate at and 90% of it at . Smaller copies use the link less and less, because each one also pays a fixed cost of a tenth of a millisecond or more (section 3).

Copy bandwidth against size, three cards; the dashed line is the link's figure

The DMA-only values at every size, per card

GB/s, the mean of five runs; each value's 99% interval is in the chart's tooltip and in pcie.json.

2. What the bounce copy costs

A program has no pinned-memory call in this runtime: every copy of user memory is staged. For a large copy the runtime fills the bounce buffer in chunks of up to with a plain memcpy on its thread pool and enables the DMA command only when all of its chunks are copied (MemcpyH2DAction.cpp in esperanto-tools-libs; a copy back runs the DMA first, then copies out), so the host copy and the DMA run one after the other. The staged rate is then close to what the two in series give, 1 / (1/DMA + 1/memcpy), with the host's memcpy measured by the same program on the same host. That is why the cards stage at different rates over the same link: the hosts' own memcpy of 256 MB runs at .

Put each card's staged rate against its host's memcpy, and every card lands on the curve of the two copies in series; the drop from the DMA-only tick above it to the point is what the bounce copy costs:

A program's copy is two copies in series: the staged rate against the host's memcpy, 256 MB

The table, card by card

3. Small copies: the runtime's polling sets the latency

A 4 KB copy spends well under a microsecond on the wire. From issue to completion, after a random idle gap, 98% of them took (every card pooled), and the spread is not noise. The runtime's response thread polls the card's completion queue; it sleeps 50 µs between polls while commands are in flight and 500 µs when none are (ResponseReceiver.cpp in esperanto-tools-libs, the constants kResponsePollingIntervalWithEventsOnFly and kResponsePollingIntervalNoEventsOnFly). A copy issued while it is in the long sleep waits for that sleep to end. The chart pools every 4 KB copy of the five runs: back to back (issued as soon as the previous one returned), or after a random idle gap of 0 to 1 ms, as a program that copies now and then would.

4 KB copies, issue to completion: how many took how long (20 µs bins)

4. Starting a kernel

The probe's kernel does nothing: each hart enters, returns 0 and goes back to the firmware. One launch, waited for, measures the whole path: the command through the submission queue, the master shire's dispatch to the compute shires, their start and return, the completion back to the host, and the runtime noticing it. A hundred launches queued at once, each with the device-side barrier, measure the card's side alone, because the host's polling then overlaps the work.

Set the two side by side, for a launch and for a small copy issued now and then, and most of the time of one operation is the wait for the runtime's response thread to wake, not the work:

Where the time of one operation goes: its cost when queued, and the wait for the response thread

The launch table

5. Several transfers at once

Each stream here moves two 64 MB copies, DMA-only. Without a barrier the runtime sends both commands at once and the card's DMA worker gives each its own channel (it has four read and four write channels, one command per channel: pcie_dma.h, dmaw.c in the firmware); with the barrier flag on every copy, one command of the stream runs at a time. The pilot run found the surprise this section is built around: two H2D commands in flight move less than one.

Aggregate bandwidth by configuration, DMA-only, three cards

Which commands collide (29 September)

A pre-registered follow-up, experiment E55 (the hub's rungs 34 and 35), asked why. Its rules were set in development on aifoundry1's card 1 (28 September, 23:47–23:55 PDT), frozen (tools/claims-v3/pcie2/PREREG.md, sha256 f632d6b3…), and validated on aifoundry3 (29 September, 00:05–00:27 PDT, five passes), then run the same way on aifoundry2 as a third card that evening (16:59–17:21 PDT, five passes, after the DVFS validation there ended). It timed two DMA-only copies per stream of 1, 4, 16 and 64 MB: both commands in flight, one at a time, and one command in each of two streams. Only two commands of one stream collide. Two host-to-card commands in flight in one stream moved 0.49 of one at 64 MB on aifoundry3 (a rate ratio of 0.48 from the fit over 1–64 MB), and 0.49 again with each command split into eight DMA elements; one command in each of two streams moved 1.01 of one. aifoundry2 gave 0.50 (a rate ratio of 0.49), 0.50 with eight elements and 1.01. A read engine or an IOMMU shared by the two commands would have halved that case too, so both are refuted, and so are a fixed cost per overlapping command, a loss that sets in only after a long overlap, and a loss per element. Development on card 1 and the run on aifoundry2 gave the same verdicts. So two streams move as much as one command alone, no more, and two commands in one stream move half as much. Card to host, two commands in flight moved 1.10 of one at 64 MB on both cards; at 1–16 MB the passes spread too wide to decide on aifoundry3 (0.78–0.94), and on aifoundry2 two 4 MB commands moved 0.88 of one, below the 0.95 predicted, a prediction no account of the collision rests on. Why one stream's two commands collide is not established.

Where a host copy lands on the card

The same experiment asked where the bytes of a host-to-card copy are written: into the L3 slice that is each line's home, or past it to DRAM. A 4 MB buffer was written by the normal (staged) copy; then one hart on shire 0 timed the first load of 4,096 of its lines, one at a time, in a kernel launched without the L3 flush. Two references were timed the same way: every line from DRAM (a median of 304 cycles) and every line from the L3 (175 cycles). A host write lands in the L3. After the copy, 99.5% of the lines read at L3 latency on aifoundry3 and 99.6% on aifoundry2 (99.9% on card 1 in development), whether or not the L3 held the buffer when the host wrote it, and every value read was the one last written. The rival accounts, an L3 that updates only lines it already holds, and writes that go to DRAM with the L3's copy invalidated or left stale, are refuted. A kernel that reads a buffer the host has just copied therefore starts from the L3, not from DRAM; that was tested for a 4 MB buffer only. Data: docs/reports/data/2026-09-29-pcie2/ (its README and results.md, every prediction with its verdict).

6. The predictions, and how they fared

The predictions were written down before any timed transfer (docs/reports/data/2026-09-27-pcie/PREREG.md, its sha256 in the data), from the link's figure, what GPUs achieve on such links, and the runtime's source. A prediction passes on a card when that card's 99% interval over its five runs lies inside the predicted range, fails when it lies wholly outside, and is inconclusive otherwise.

7. What this does and does not show

Method, and how to reproduce it

The probe. workloads/pciebench is a standalone host program and an empty device kernel in the et-testdrive style, using only the runtime API. One test per process, each under timeout 10 with an 8.5 s budget and the card's lock held for the whole run: info (device properties, DMA limits, the negotiated link from sysfs), bw (the sweep, with a 1 MB there-and-back data check on the staged path, which passed in every run: ), lat (64 B and 4 KB round trips back to back and after random gaps, 200 pipelined 4 KB copies, an idle-stream wait), launch (150 single launches and three batches of 100 per shire mask, twice), conc (eight configurations, three trials) and hostcopy (the host's memcpy, no card). A telemetry sample before and after each run records the die temperature and clocks.

What the pilots changed. The first pilot on aifoundry2 ran the four copy variants in a fixed order, and each variant locked onto one of the runtime's two polling modes (one variant always about 0.11 ms, the next always about 0.58 ms at the same size: the round-trip medians in raw/aifoundry2/pilot1/lat.out). The schedule's probe shuffles the order in every repeat of the bandwidth sweep, and adds 4 KB copies after random gaps, also shuffled. Its back-to-back round trips (section 3's table, and predictions P6 and P7) kept the fixed order, host to card staged, then DMA-only, then card to host staged and DMA-only (workloads/pciebench/host/main.cpp, testLat), so they can still lock each variant onto one mode. The same pilot showed two queued H2D commands at half the rate of one, so the concurrency test gained the barrier configurations (the /ser rows). Neither change touched a prediction.

Reduction. workloads/pciebench/reduce_pcie.py: a run's value for a cell is the median of its repeats, a card's value the mean over its five runs with a 99% t-interval (t = 4.604 for five runs).

cmake -S workloads/pciebench -B build/pciebench -DCMAKE_PREFIX_PATH=/opt/et -Wno-dev && nice cmake --build build/pciebench -j4
TREE=$PWD workloads/pciebench/run_pcie.sh 1            # one run on this host's card (V3_DEVICE=1 on aifoundry1)
workloads/pciebench/schedule.sh 1 5                     # from aifoundry2: five rounds on the three cards
python3 workloads/pciebench/reduce_pcie.py docs/reports/data/2026-09-27-pcie/raw \
    --prereg docs/reports/data/2026-09-27-pcie/PREREG.md \
    --out docs/reports/data/2026-09-27-pcie/pcie.json --md docs/reports/data/2026-09-27-pcie/results.md
python3 scripts/build-report.py pcie-link docs/reports/data/2026-09-27-pcie/pcie.json \
    docs/reports/2026-09-27-et-soc1-pcie-link.html

Data: docs/reports/data/2026-09-27-pcie/ (its README lists every file).

Version history and provenance

First written 27 September 2026, from the runs of that afternoon; later that day, the staged path drawn against the hosts' memcpy (section 2) and the time of one launch or small copy split into its queued cost and the wait for the response thread (section 4); 28 September, the hosts' memory channels (section 2), read as a user the night before, where the page had said reading them needs root; 29 September, which commands collide and where a host copy lands (section 5), from E55; the page had left the first open. 30 September: E55's third card, aifoundry2, in section 5. Every number on this page is computed by its script from docs/reports/data/2026-09-27-pcie/pcie.json, except section 5's two follow-ups, whose figures are quoted from docs/reports/data/2026-09-29-pcie2/results.md (aifoundry3 and aifoundry2) and dev-aifoundry1-c1.md (card 1, development), and the link figure's arithmetic (16 GT/s × 8 lanes × 128/130 ÷ 8, also in the data) and the runtime's and firmware's constants, which are cited to their source files above.