Summary
- CAIDA’s integrated active-measurement environment places reusable measurement primitives between unrestricted access to a vantage point and a results-only service. That design can make the kind of traffic a platform permits easier to explain to the organizations hosting its probes.
- The access relationship has three parties: the site that hosts a vantage point, the platform operator that maintains it, and the researcher who uses it. Each has a different authority and bears a different kind of risk.
- Matthew Luckie is an important author of Scamper and the 2025 programming-environment work, but the public record is collaborative. A constrained interface can reduce ambiguity; it cannot certify that every measurement is authorized, harmless or representative.
The host behind the probe
An Internet measurement begins somewhere. Before a route, delay, DNS response or web endpoint appears in a paper, a packet has left a network from a particular vantage point. The organization hosting that point may be a university, a research institution or another site that has agreed to provide connectivity and equipment. Its first operational question is concrete: what will the machine in our network send, to which destinations, how often, and under whose control?
That question is easy to miss when measurement is described only as a scientific method. A probe is not just an observation. It is traffic carried by a real network, scheduled by software and delivered to a recipient that may not be participating in the experiment. Its source address can be seen by routers, firewalls, destination services and abuse desks. Even when the researcher intends only to measure, the host site's network carries the traffic and can be associated with it.
Matthew Luckie’s work at CAIDA is useful because it makes this operational boundary part of the design problem. He is the author of Scamper, a packet prober used by CAIDA’s Archipelago, or Ark, infrastructure. In later work with a team of coauthors, he helped present a programming environment that exposes measurement functions as reusable building blocks. The point is not that this interface settles every question about a probe. It is that the platform can describe a bounded set of things a researcher may ask a vantage point to do, instead of treating the vantage point as an open-ended machine. Luckie’s CAIDA profile and the team’s 2025 PAM paper frame the work in those terms.
Three parties, not one authorization
The 2025 paper names three actors: the site that hosts a vantage point, the operator that runs the measurement platform, and the researcher who uses it. Their positions are related but not interchangeable. A site supplies the location and network attachment. The platform operator maintains the system and decides what capabilities it exposes. The researcher defines an experiment and receives data from the measurements. The site incurs risk by hosting a point and must trust the platform operator not to use it in ways that harm the host network. The platform operator, in turn, takes on risk when it grants researcher access.
This division matters because “permission” can refer to different decisions. A researcher may be approved to use an academic platform. That does not by itself show that a specific destination consented to receive a probe. A host may agree to run a measurement node. That does not establish that the host reviewed every later experiment or its precise targets. A platform may constrain the kinds of packets a researcher can request. That does not prove that the constraint is sufficient for every target, rate or purpose.
The evidence available for Ark illustrates a layered rather than universal approval process. CAIDA’s current Ark programming guide says that vetted academic researchers may receive access to a CAIDA system for on-demand measurement from Ark vantage points. The public access request form asks for an experiment’s goal and expected load, including measurement types, destinations, number of measurements, frequency, project duration and vantage-point requirements. It includes an acceptable-use agreement and requires publication reporting. These are observable parts of an admission process. They do not reveal whether a particular host separately approves every target or whether every approved experiment has the same risk profile.
The current guide adds a concrete host-facing control: measurement capabilities differ by vantage point according to the host’s preference. CAIDA says every listed Ark point supports ping and traceroute, most also support DNS, UDP and HTTP, and a small number support OWAMP. The Python module exposes tags so researchers can inspect a point’s available capabilities before scheduling an operation. This makes host preference visible in the usable capability set; it still does not establish target-specific consent or individual host review of every later experiment.
Treating the three roles separately also prevents a common analytical shortcut: interpreting participation as a mandate. A university or network that hosts a measurement point is an affected stakeholder and an operational participant. The measurement platform is not thereby authorized to speak for every destination network. The researcher’s approved account does not speak for an unconsulted recipient. The host, platform and experimenter need distinct records of what each has agreed to do.
The middle of an access spectrum
Measurement systems distribute capability in different ways. At one end, a researcher can receive direct shell access or run code on the vantage point. That offers flexibility but gives the platform operator fewer ways to tell a host what the code may do. A sandbox can limit resources or packet patterns without necessarily making every experiment's effects obvious in advance. At a more restrictive end, a service may allow only a fixed set of tests or publish data already collected by its operator. That gives the user less control over methods and timing.
The integrated environment described by Luckie and his coauthors sits between those endpoints. The paper's design exposes a collection of measurement primitives through a Python interface. A user can assemble a more involved workflow from operations such as ping, traceroute, DNS, HTTP, UDP probing, alias resolution and TCP-behaviour tests. The platform can explain the exposed capabilities to a hosting site, while researchers can combine them without implementing every packet-handling detail themselves.
In the paper's account, the interface is designed to retain useful programmability while making the traffic classes more describable than an unrestricted packet-sending socket would be.
That is not the same as granting arbitrary code execution on each Ark node. The paper separates measurement logic implemented at the vantage points from experiment logic coordinated at or near a central controller. The current Ark programming documentation describes access to a CAIDA system, and Luckie’s 2024 explanation of the Scamper Python module says that access is to a system that can invoke measurement primitives, not a login on each Ark node. The distinction matters to hosts: which executable runs locally, which decisions remain centralized, and which packet types may leave the site are separate questions.
It also matters to researchers. A restricted API can be too narrow for a new experiment; unrestricted code can be too broad for a site to accept. The middle model tries to make the useful part explicit: let a researcher express measurement intent using operations whose implementation and effects the platform can document. The resulting compromise can widen access for a researcher who does not want to maintain low-level instrumentation, while retaining an operator-controlled capability catalogue.
Why primitives change the work
The interface is not only a permission layer. It addresses a long-standing engineering cost in Internet measurement: getting packet behaviour right across systems, methods and response types. In his 2010 paper, “Scamper: a Scalable and Extensible Packet Prober for Active Measurement of the Internet”, Luckie describes the need for software that can execute measurements consistently and systematically. The tool implements techniques such as traceroute, ping, MDA traceroute and alias resolution so researchers can focus more of their work on the experiment than on rebuilding instrumentation.
The distinction between an instrument and a research question is consequential. A traceroute variant can send packets differently from a system’s built-in tool. A measurement may require several protocol types, carefully timed probes, or observations from multiple vantage points. If every researcher has to reimplement the low-level mechanics, inconsistent implementations can become a source of experimental variance. A shared prober can make the method easier to repeat and the code easier to inspect. But standardization at the instrument layer does not remove assumptions in target selection, sampling, interpretation or analysis.
The 2025 paper describes a Python layer over Scamper that normalizes an uneven set of lower-level interfaces. The ScamperCtrl abstraction can coordinate several vantage points, schedule a mix of synchronous and asynchronous measurements, and return results through common methods. The authors report that their Cython bindings span roughly 11,000 lines. Those details explain why an interface can change who is able to write an experiment: users can work with a familiar language and focus on how a sequence of measurements answers a question, rather than learning every controller detail first.
They also explain why a primitive list has power. If researchers can invoke DNS, HTTP, alias resolution, UDP probes or TCP-behaviour tests through a common controller, the platform has chosen and implemented a meaningful set of actions. Adding a primitive can expand what an experiment can observe. Removing or changing one can break old code or make a previously possible question harder to ask. The operator's catalogue is therefore not neutral plumbing. It is a technical policy surface, even if the public paper presents it primarily as an engineering design.
A host can be told more, not everything
The key design argument in the 2025 paper is host transparency. A measurement primitive can be described in terms that a platform operator can communicate to the site hosting a vantage point. A raw packet interface offers a weaker advance description: a user may construct packet sequences that the host cannot easily infer from the name of a generic permission. A primitive such as a DNS query, an HTTP request or a traceroute narrows the vocabulary of possible actions.
That narrowing can make a promise more credible. If the published system says that users can request specific measurement functions, and the platform implements those functions rather than offering arbitrary packet injection, a host has a clearer basis for evaluating what the machine is intended to do. A well-documented primitive can state supported protocols, parameters and operational constraints. It can also let an investigator read a research script and see which operations are called. This is a material improvement over a vague assurance that “researchers will be careful.”
But a vocabulary is not a complete policy. A named HTTP primitive does not, by itself, say which destination is appropriate, how many requests are acceptable, which content will be fetched, what a recipient will infer from the traffic or how an experiment will react to errors. A DNS lookup can be innocuous in one context and part of a large-scale enumeration in another. A traceroute can help diagnose a path, yet repeated probing of a sensitive endpoint may still cause operational concern. The relevant unit of risk is not only the primitive; it is the combination of capability, target, rate, duration, purpose and the recipient's environment.
The platform operator can put limits in code, choose available methods, approve accounts and communicate expected traffic. The researcher remains responsible for choosing the target list and designing a scientifically defensible experiment. The site host retains its own operational exposure and may require clear procedures for escalation or suspension. The public sources establish that CAIDA has developed an access environment and requires vetted researchers to submit experiment details. They do not establish a universal per-target consent rule or an independent audit of every safety control.
Those should remain open questions rather than gaps to be filled with reassuring language.
Build locally, run through a shared controller
The development path described by Luckie shows why usability and access control are linked. In a January 2024 CAIDA blog post, he describes testing Scamper code on a workstation or local Raspberry Pi, then using available Ark vantage points after moving the experiment to the CAIDA system. The earlier DSL post describes the Python module as an interface to measurement primitives on Ark nodes and says researchers can develop and test complex measurements locally before running vetted experiments.
This division can lower the cost of experimentation. Researchers can debug basic logic without consuming remote vantage-point resources, and they do not have to reproduce every feature of Scamper at each site. The platform can centrally coordinate requests while the distributed nodes perform the defined measurement operations. In the paper's architecture, a controller mediates between the Python environment and Scamper processes at the vantage points. It allows code near the controller to compose a distributed experiment without asking every host to administer accounts for individual researchers.
The same feature concentrates responsibility. If the controller is the boundary through which experiments are scheduled, then its authentication, logging, quota policy, failure handling and ability to stop work matter. A site may rely on the platform to honour the description it supplied; researchers may rely on the controller to return results from the vantage points they selected. Those obligations are not eliminated by moving code out of the node. They move to a different layer and must remain visible.
The current Ark form makes several parts of the request legible before access is granted. It asks where the user will measure, how targets are selected, how many probes will be sent, how often, for how long, and whether particular vantage-point properties are required. If the request is approved, CAIDA asks for the user's SSH public key. That process does not prove how every host is consulted, how platform-side limits are implemented or what happens after a target complains. It does show that access is presented as an experiment-specific request rather than an anonymous public endpoint.
Composability in practice
The paper's examples help show what the programming environment makes possible, and what it does not claim. A short script can issue ping measurements from a set of vantage points and select the smallest observed round-trip time. A different script can find the authoritative name servers for a domain, resolve their addresses and measure delay to each. These are not only convenient wrappers: the experiment contains dependent steps, so the interface must make sequencing, parallel work and partial failures manageable.
A more involved example follows the server selection behind Netflix's Fast.com service. The researchers combine DNS lookups, HTTP requests and traceroutes from different Ark vantage points to identify the speed-test servers returned to each location and then observe path delay. In a four-day May 2024 illustration from a vantage point in Thimphu, the paper reports that latency to servers in Hong Kong or Singapore occasionally rose substantially and that Fast.com returned US servers during some of those intervals. The researchers describe traffic load as appearing to influence the selection.
This is a bounded case study from a particular vantage point and period, not evidence of a general Netflix rule or a ranking of service quality.
The authors also implemented components associated with MIDAR, a technique that uses observations from multiple vantage points to infer when different IP addresses may correspond to interfaces on one router. They report replacing 2,554 lines of Ruby with a 902-line Python script for one workflow. The reduced script can make coordination logic easier to read, but shorter code is not automatically more correct. Alias inference depends on how probes are scheduled, which responses arrive and what assumptions connect an IP-ID pattern to a common device.
The source paper provides one example of implementation and scale; it does not turn every inferred node into directly observed physical equipment.
These cases support a modest claim: composable primitives can help researchers express multi-stage and distributed experiments without hand-building as much controller code. The examples do not establish that every user can run every script, that every destination welcomes the resulting traffic, or that the observation is true outside the sampled paths. Operational ease and evidentiary reach are different variables.
Counts need a date and denominator
The scale of a measurement network can tempt readers to treat its data as complete. The PAM paper describes Ark as having roughly 170 vantage points in 57 countries and 133 autonomous systems in October 2024. Those values are a snapshot explicitly dated by the authors, not a current count for 2026. They indicate that Ark can offer a distributed set of observation points; they do not mean that every country, network type, access technology or routing relationship is represented.
CAIDA’s 2025 annual report later says Ark expanded to approximately 300 active vantage points in 2025. The paper’s October 2024 figure and the report’s 2025 estimate are separate snapshots with different dates and wording; without a shared definition and counting method, they should not be treated as a directly comparable growth series.
The paper's comparison of Internet Topology Data Kits from February 2023 and February 2024 illustrates why denominator discipline matters. It reports that the number of vantage points with traceroute data rose from 93 to 142, and the number of countries from 37 to 52. The probed address count increased from 2.64 million to 3.58 million. The authors say the growth was driven by expansion of Ark vantage points. A larger observation set can reveal more router-interface addresses, but the count of observed addresses is not the count of Internet routers.
The authors explicitly call their inferred multi-address graph nodes “nodes” to distinguish inference from direct identification of physical routers.
That language is more than a technical caveat. A path observed from one vantage point may differ from a path observed from another. A returned source address may belong to an interface not on the forward path. A missing response can be caused by filtering, loss, rate-limiting or a measurement limitation. A measurement primitive standardizes how a probe is emitted; it does not standardize the Internet's response or guarantee that the sample is balanced.
For the same reason, the paper's selected Netflix example should not be read as a complete map of CDN behavior. The displayed time window, vantage, request logic and returned endpoint list define what is observable. A pattern in one sequence can suggest a hypothesis about latency and selection. A broader claim requires new observations and an explicit method. A reusable programming environment may make that next experiment easier to build; it cannot substitute for one.
What Luckie's record supports
Matthew Luckie's public profile lists Scamper and the 2025 integrated programming environment alongside work on routing, topology and measurement. That record supports writing about his contribution to measurement tooling and the boundary between exposed capability and host expectations. It does not support assigning him sole authorship of the environment or sole authority over Ark's admission decisions.
The 2025 paper has seven authors: Luckie, Shivani Hariprasad, Raffaele Sommese, Brendon Jones, Ken Keys, Ricky Mok and k claffy. Its acknowledgements say Bill Herrin suggested that a domain-specific language could accelerate active-measurement discovery at a CAIDA AIMS workshop, and Alexander Marder suggested beginning with Python bindings for Scamper. That lineage is important. A first author can be a central contributor without being the only person who shaped the idea, wrote the software or manages the institution around it.
The public record also places the programming environment inside an ongoing CAIDA research programme. The current Ark guide lists experiments involving public resolver behaviour, anycast, MPLS tunnels, DNS deployment and other topics. The 2025 CAIDA annual report describes the environment as a way to lower the barrier for researchers while helping operators bound the measurements users may run. These are CAIDA's own statements about a system it develops and operates. They are useful evidence of intended design and current institutional presentation, not independent proof that every promise has been met at every host.
For an article about a person, this distinction guards against two kinds of error. The first is hero-making: presenting a collaborative platform as one individual's invention. The second is false institutional attribution: treating the named author as if he alone approves access, represents every site host or controls all uses of the infrastructure. The public material documents Luckie's technical role and his writing about Scamper. It does not expose the full internal division of labour for every operational decision.
A useful boundary, not a certificate
An integrated measurement interface can improve three things at once. It can reduce the programming work needed to assemble repeated operations. It can make a platform's available capabilities more specific than “send packets from here.” And it can help the platform operator explain to a hosting institution what its measurement nodes are intended to do. Those are meaningful benefits in a field where researchers need observations from networks they do not own.
Each benefit has a limit. A named primitive remains only as safe as its implementation, parameters and schedule. A host-facing description can be clear yet incomplete. A vetted research account can still select poor targets or run an overly aggressive experiment. A correct measurement from a limited collection of vantage points can still be unrepresentative. None of these observations makes the CAIDA design ineffective; they mark the part of the problem that the design does not purport to solve.
The strongest conclusion is therefore architectural rather than personal. Luckie's Scamper work and the team's 2025 environment move the boundary of active measurement from an implicit assumption about user behaviour toward an explicit catalogue of capabilities and a more legible access relationship. The remaining work is institutional: keep that catalogue aligned with deployed primitives, state the limits and rates that hosts are actually relying on, preserve accountable experiment descriptions, and provide a way to respond when the traffic or inference exceeds the promise.
The probe has a host. A measurement platform earns continued access not by declaring its probes safe, but by making the capability, purpose, scope and responsibility behind them inspectable—and by leaving room for the hosting network to challenge or withdraw from the arrangement.
Sources
- Matthew Luckie — CAIDA profile
- Scamper: a Scalable and Extensible Packet Prober for Active Measurement of the Internet (IMC 2010)
- An Integrated Active Measurement Programming Environment (PAM 2025)
- Ark programming environment — CAIDA
- Ark access request — CAIDA
- CAIDA 2025 annual report
- Towards a Domain Specific Language for Internet Active Measurement — CAIDA blog
- Developing active Internet measurement software locally to run on Ark — CAIDA blog
- Scamper software catalogue — CAIDA
- Scamper Python module documentation — CAIDA
- CAIDA 2025 annual report
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
