Repository navigation
Conversation
The deployment switch has PTP enabled but its grandmaster has no GPS, so
its absolute time is not a usable reference. The transmitter must own the
reference clock instead.
MTL's built-in PTP client is slave-only -- mt_ptp.c only ever transmits
PTP_DELAY_REQ and never sends Announce/Sync -- so MTL cannot elect itself
grandmaster. The role is therefore split: ptp4l serves the NIC's PTP
Hardware Clock to the receivers, and dvledtx feeds MTL that same PHC
through mtl_init_params.ptp_get_time_fn with pacing forced to
ST21_TX_PACING_WAY_PTP. TX pacing, RTP timestamps and the clock the
receivers lock to then all resolve to one oscillator, independent of the
switch.
Add a ptp.mode option selecting between the existing external-grandmaster
behaviour ("slave") and the new local-PHC reference ("grandmaster"),
alongside phc_interface / phc_device / phc_interval_ms.
New util/ptp_clock serves the PHC to MTL. A /dev/ptpN read costs
microseconds, far too slow for the per-packet pacing path, so the PHC is
sampled on a background thread and served as a frequency-corrected linear
model over CLOCK_MONOTONIC_RAW, published through a lock-free slot ring.
PHC references from config are validated against /dev/ptpN or an interface
name before reaching open().
Grandmaster mode requires the direct MTL TX build: a function pointer has
no AVOption equivalent and the mtl_st20p plugin owns the mtl_init() call,
so the muxer build rejects it at config load rather than transmitting on
the wrong clock.
Also fix a latent bug on the muxer path: ptp.enable set only the
MTL_FLAG_PTP_* flags and never pacing_way, leaving mtl_init_params.pacing
at AUTO, which resolves to TSC. Slave mode therefore never actually used
PTP pacing there.
Add scripts/ptp_grandmaster.sh (ptp4l with masterOnly 1 and priority1 64,
so BMCA stays on this host against switch grandmasters that advertise
128), a sample config, a README section and unit tests.
Verified: both build variants compile warning-clean; the PHC time source
tracks a live Intel E610 PHC to a 256 ns residual with a 38930 ppb
frequency correction.
…eiling Found while bringing the feature up end to end on an E610. scripts/ptp_grandmaster.sh: - linuxptp 4.0 renamed masterOnly to serverOnly and warns on the old spelling. Detect the major version from `ptp4l -v` and emit the right key, keeping masterOnly for linuxptp 3.x. - The script ends in `exec ptp4l`, which replaces the shell, so the `trap ... EXIT` that was meant to clean up the mktemp config could never fire and every invocation leaked a file in /tmp. Write a fixed /tmp/ptp4l-gm-<iface>.conf that overwrites instead of accumulating. README: document that sessions per VF are bounded by the NIC's TX queues. MTL requests sessions+2, silently clamps to what the device provides, and only fails later with "fail to find free tx queue". An E610 VF has 4 TX queues and MTL reserves one, so 3 sessions per VF is the ceiling. Add two configs covering both sides of that limit: 3 sessions on a single VF (at the ceiling) and 4 sessions spread one-per-VF (scaling past it). Validated on an E610: all sessions report "pacing way: ptp", the PHC servo settles to tens of ns with a steady ~40 ppm correction, and frame get/succ/put match with drop 0.
DPDK's ixgbe_timesync_enable only installs an ethertype 0x88F7 filter, with no UDP-319 timestamping. ptp4l defaults to network_transport UDPv4, so an MTL receiver could never hardware-timestamp our Sync messages. Verified on an E610: over UDPv4 the receiver's PTP never locks; on L2 it reaches SLAVE state and converges to 31 ns rms.
aa56dff added network_transport L2 to the script but documented it nowhere, so anyone writing their own ptp4l config would hit the failure it fixes: DPDK's ixgbe timesync installs only an ethertype 0x88F7 filter and never timestamps PTP over UDP, so a UDPv4 grandmaster runs fine locally while no MTL receiver can ever lock to it. Also note that the script emits serverOnly on linuxptp 4.0+, matching the version detection already in the script.
… tests The unit test job failed to build. mtl_tx.c now calls ptp_clock_open, ptp_clock_get_time_ns, ptp_clock_device and ptp_clock_close, but ptp_clock.c was only added to the test_config_reader target, so test_mtl_tx, test_mtl_tx_stub and test_session_manager all failed to link with undefined references. scripts/test.sh sends meson and ninja output to /dev/null, so CI showed only "Setting up test build..." followed by exit 1 with no diagnostic. Extend the PTP coverage while here: - explicit mode "slave", and a negative phc_interval_ms clamping to 0 - empty phc reference accepted, since that selects the host's only PHC - interval boundaries 100/10000 accepted and 99/10001 rejected - slave mode ignores the phc fields rather than validating them - interface name length boundary at IFNAMSIZ - ptp_clock_open on a missing device fails, and a NULL handle is safe through every entry point 151 tests pass locally against cmocka 1.1.7.
ptp_clock.c sat at 20% line coverage: the config-level tests reached validation and the NULL paths, but the servo that the whole feature depends on was untested. sampler_update() already takes the two timestamps as plain integers -- the syscalls live in its callers -- so the servo can be driven with synthetic samples and no PHC hardware. The test includes ptp_clock.c directly to reach the static functions, as test_main.c does for main.c, and wraps clock_gettime so the served time is deterministic. Covers first-sample anchoring, convergence on a 40 ppm drift, the step-on-jump threshold, the stalled-clock guard against a divide by zero, rate clamping, continuity across a re-anchor, monotonicity under noisy samples, and the slot ring wrapping. Verified non-vacuous by mutation: re-anchoring on the raw measurement instead of the estimate, making clamp_rate a no-op, and raising the step threshold each fail the suite. The first revision of the clamp test did not catch its mutant -- a 1 s discrepancy takes the step path and never reaches the rate calculation -- so it now uses 0.9 ms over 100 ms, which stays under the step threshold and implies +9000 ppm.
The README told users to spread sessions one-per-VF "as in config/tx_ptp_gm_4session.json". That config cannot work: MTL segfaults on any port after the first once PTP frames arrive, so following the advice crashed immediately. mt_ptp_init() allocates a PTP instance only for port 0 unless a port reports offload timestamping, which ixgbe VFs do not, leaving impl->ptp[i] NULL for ports 1+. cni_rx_handle() then calls mt_ptp_parse() for ethertype 0x88F7 with no NULL check, though the UDP path just below it is guarded. ptp4l runs on the PF and the NIC's internal switch replicates its L2 multicast to every VF, so the second VF dies on the first PTP frame. Measured on an E610: 3 sessions on 1 VF runs clean, while 4 sessions on 2 VFs and 8 on 4 VFs both SIGSEGV in mt_ptp_parse. It reproduces with ptp.enable false and stops only when no PTP traffic is present, so it is independent of this feature. Re-confirmed against a DPDK built with MTL's documented patches. Remove the config, document the real ceiling of 3 sessions on a single VF, and record the one-line mt_cni.c guard that lifts it. Add a startup warning for multi-interface grandmaster configs so the failure is explained rather than a bare segfault.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds hardware PTP timing to
dvledtx, including agrandmastermode that lets the transmitter actas the reference clock for the media LAN.
The motivating deployment has a PTP-capable switch, but its grandmaster has no GPS and therefore no
true time, so it cannot serve as the reference. The transmitter's NIC PHC is used instead. Absolute
time does not matter here — only that every receiver inherits the same clock.
Design
MTL cannot be a PTP master:
lib/src/mt_ptp.conly ever sendsPTP_DELAY_REQand never emitsAnnounce or Sync. The grandmaster role is therefore split in two:
ptp4lserves the NIC's PHC to the network (scripts/ptp_grandmaster.sh)dvledtxfeeds MTL that same PHC viamtl_init_params.ptp_get_time_fn, withpacing = ST21_TX_PACING_WAY_PTPBoth therefore derive from one clock.
MTL_FLAG_PTP_ENABLEis deliberately not set, so MTL's ownclient stays out of the way.
src/util/ptp_clock.csamples/dev/ptpNon a background thread (1 s default) and serves MTL afrequency-corrected linear model over
CLOCK_MONOTONIC_RAWthrough a lock-free slot ring. A raw PHCread costs microseconds — far too slow to do per packet on the pacing path.
No changes to MTL, DPDK or FFmpeg are required. This uses existing public API only, and the PHC
is read through the kernel rather than DPDK.
Notable fixes included
ffmpeg_tx.csetptp.enableflags but left pacing atAUTO, which silently resolved to TSC.It now sets the
pacing_wayAVOption, so the muxer path actually paces to PTP.scripts/ptp_grandmaster.shforcesnetwork_transport L2. This is required, not cosmetic: DPDK'sixgbetimestamping installs only an ethertype0x88F7filter and never timestamps PTP over UDP,so a UDPv4 grandmaster runs fine locally while no MTL receiver can ever lock to it.
ptp4lconfig used a path thatexecmade unreachable from theEXITtrap, andthe
masterOnlykey was renamedserverOnlyin linuxptp 4.0. Both handled.Validation
Hardware-tested on an Intel E610:
pacing way: ptp; MTL logsuse user ptp source, confirming the callback isin use rather than TSC or
CLOCK_REALTIME.drop 0, ~435 Mb/s/session.(
selected local clock f80278.fffe.2491bb as best master).ptp4lreachesSLAVE on MASTER_CLOCK_SELECTEDwithrmssettlingto ~31 ns, and the receiver adopts the transmitter's time.
Unit tests for the new config parsing are added in
tests/test_config_reader.c.Reviewer notes
cmockais absent on the test host, sotests/meson.buildsilentlysubdir_done()s. They need a run on CI or a host with cmocka.ceiling. MTL requests
sessions + 2, silently clamps, then fails later with the unhelpfulmt_dev_get_tx_queue(0), fail to find free tx queue. Documented in the README, andconfig/tx_ptp_gm_4session.jsonspreads sessions across VFs to stay under it.grandmastermode requires the PHC interface to stay kernel-bound, so it needs either SR-IOV or aseparate PTP port. Both topologies are documented.
grandmastermode on the FFmpeg-muxer build (#ifndef ENABLE_MTL_TX).