Summary
- Candela’s documented work spans BGPlay, RIPE Atlas streaming and visualisation, DNSMON, TraceMON, RIPE IPmap, geofeed monitoring and the open-source BGPalerter project.
- These systems compress routing updates, traceroutes, latency samples and validation records into timelines, enriched paths and alerts that can shorten the route from signal to investigation.
- Every result remains viewpoint-bound: collector coverage, probe placement, feed freshness, enrichment methods and threshold choices decide whether a change describes one observer or a broader incident.
- His move from research to RIPE NCC public infrastructure and NTT operations shows a consistent discipline at different scales, while current ownership and maintenance must be stated tool by tool.
BGPlay turned route change into a sequence an operator could replay
Around 2012, Massimo Candela’s work on BGPlay, later integrated into RIPEstat, turned a BGP incident from a file of updates into a sequence an operator could replay. A prefix could disappear, return with another origin, become more specific or converge differently across collectors. The interface reconstructed an initial view, applied announcements and withdrawals in time order and made the transition visible as a changing graph.
That representation solved a cognitive problem rather than a routing one. A raw feed contains more detail than a diagram, yet the detail can hide the moment that matters. BGPlay selected, grouped and ordered observations so a user could ask when an origin appeared, which paths changed and which collectors saw the event. It also inherited every limit of the underlying evidence: collector coverage, AS-level aggregation, delayed updates and layouts that can make one relationship look more central than it is.
Candela carried the same method from research into RIPE NCC public measurement systems and later into continuous monitoring. RIPE Atlas interfaces, DNSMON, TraceMON and RIPE IPmap combined measurements with context; BGPalerter changed the interface from an investigation opened by a user into an alert that interrupts one. Geofeed work gave operators another way to publish structured claims.
The governing question is how an interface can shorten the path from distributed signal to operational judgement without converting partial evidence into certainty. The answer depends as much on provenance, thresholds, maintenance ownership and notification delivery as on visual design. Candela’s record is strongest where the tool makes the next question easier and leaves the underlying observation available for challenge.
BGPlay made route evolution navigable without claiming topology ground truth
This model is particularly useful when several changes overlap. A legitimate origin may withdraw before an unexpected origin appears. A more-specific route may attract traffic even while the covering route remains. Several collectors may converge at different times. The sequence can help distinguish a brief propagation artefact from a sustained change.
BGPlay cannot show physical forwarding with router-level precision. BGP paths are control-plane advertisements at the AS level. They do not reveal internal routing, private interconnections invisible to the collector, MPLS paths or every alternative available inside a network. The visualisation should be read as what selected observers learned, not as a map of where each packet travelled.
The design also embodies choices about aggregation. Repeated updates can be compressed. Similar paths can be grouped. Labels and colour can emphasise origin or event type. Those choices make the tool usable and can hide churn or uncertainty. An expert interface needs a route from the summary back to the underlying record.
BGPlay’s institutional path matters to Candela’s profile. A research prototype became part of a public service maintained by the RIPE NCC. That transition imposed requirements beyond publication: API integration, service continuity, documentation, browser performance and support for users with different levels of routing knowledge.
The service is not personally owned by Candela. His contribution includes design and development, while RIPE NCC teams operate the institutional platform and later maintain the code. This distinction between authorship and current responsibility will recur throughout his career, most visibly in RIPE IPmap.
RIPE NCC turned interface design into public measurement infrastructure
Candela joined the RIPE NCC in August 2013 as a senior software engineer in research and development. The organisation operates RIPE Atlas, RIPE RIS, RIPEstat and related services used by networks and researchers. Working inside that environment changed the scale and lifecycle of his projects.
RIPE RIS collects BGP information from routing peers at distributed collectors. RIPE Atlas uses a global network of probes and anchors to perform active measurements such as ping, traceroute and DNS queries. RIPEstat provides interfaces to internet-number and routing data. These systems produce different evidence and share a challenge: the raw volume and distribution make manual interpretation impractical.
Candela’s RIPE work concentrated on interfaces and streaming systems that let users turn the platforms toward a focused question. A measurement service becomes more valuable when an operator can move from “the data exists” to “these probes saw the delay change at this time” or “these collectors observed the origin transition.”
Institutional operation adds constraints that research prototypes can defer. Public APIs need versioning. Live streams can be incomplete or delayed. Visual tools have to handle users who do not understand every data-quality caveat. Services need monitoring, security and maintenance after the original developer leaves. Method changes must be documented because people may compare results across years.
The public nature of RIPE data also creates an accountability advantage. Users can often inspect the measurement identifier, probe list or routing source behind an interface. That makes it possible to reproduce or challenge an interpretation. The platform remains partial, but its partiality can be described.
The RIPE period gave Candela a broad portfolio: BGPlay and RIPEstat work, RIPE Atlas streaming and visualisation, DNSMON, LatencyMON, TraceMON and RIPE IPmap. These were not one product family designed by one person. They were institutional services with distinct teams and purposes. The common thread is the attempt to make distributed measurements useful to an operator before the details become overwhelming.
Streaming measurements reduce delay and create ordering problems
A conventional measurement workflow submits a task, waits for completion and downloads the stored result. That model is suitable for many studies and slow for incident response. RIPE Atlas streaming allowed applications to receive results as probes produced them, enabling interfaces to update while a measurement was still running.
Candela worked on systems that made these streams usable in web applications and operational tools. Live data can reveal the first signs of a reachability or latency change without waiting for the full campaign. An operator can see whether the problem is concentrated in a region, probe group or network and decide where to investigate.
Streaming does not convert distributed measurement into a perfectly ordered feed. Probes have different connectivity and clocks. Results can arrive late, be retried or fail. A live consumer may see a partial picture that changes when stored data is reconciled. The interface should communicate that incompleteness rather than presenting each early result as final.
Measurement identifiers and probe metadata are therefore essential. A value without the identity and status of the probe is weak evidence. A probe behind a home router, an anchor in a data centre and a device with intermittent connectivity do not carry the same operational meaning. Users need filters and enough context to avoid treating every sample equally.
Live visualisation also creates a temptation to optimise for movement. An animated graph feels responsive even when the underlying change is noise. Candela’s work is strongest when the interface directs attention toward a testable hypothesis and preserves the ability to inspect the data, rather than turning measurement into spectacle.
The streaming model later reappeared in BGPalerter, although the product form changed. Instead of waiting for a user to open a tool, the system continuously consumes routing feeds and sends a notification when configured conditions are met. The move from interactive exploration to automatic alerting increased the need for explicit rules, reliable delivery and change context.
DNSMON and LatencyMON applied active measurement to service behaviour
Routing visibility is only one layer of internet operation. A route can be present while a service is slow or failing. DNSMON used RIPE Atlas measurements to help operators inspect the performance and reachability of DNS root and top-level-domain infrastructure. LatencyMON provided ways to compare delay over time.
Active measurements ask a controlled question from selected vantage points. A DNS query can show whether a resolver reaches an authoritative server and how long the exchange takes. A ping can reveal round-trip delay and loss. Repeating the measurement across probes and time creates a view of geographic and network variation.
The strength is direct service evidence. A BGP announcement says a path exists in the control plane; an active query tests whether a protocol exchange succeeds from a probe. The weakness is coverage. A probe network is uneven, and the result describes the path between that probe and the target. Users elsewhere may have a different experience.
Latency also requires statistical interpretation. A single high sample can reflect congestion, probe load, queueing or a transient path. Medians, distributions and baselines are more useful than isolated values. The interface has to show change without disguising natural variability as an outage.
DNS adds caching and anycast. A root or TLD service can be announced from many sites under the same address. The probe’s path and resolver behaviour determine which instance is reached. A performance shift may result from routing, a server issue or a change in local resolver state. Active measurement narrows the possibilities; it rarely identifies the cause alone.
Candela’s contribution in these tools was to build interpretation layers around RIPE Atlas evidence. They broaden the profile beyond BGP and show a consistent method: collect a distributed signal, attach time and context, and let the user compare views while keeping the measurement boundary visible.
TraceMON enriched traceroutes without pretending every hop was known
Traceroute lists responding addresses along a path, subject to the behaviour of routers, load balancing, ICMP filtering and tunnels. Raw output can be difficult to interpret. An address may belong to an interface whose role is unclear. Several hops may sit inside one autonomous system. An internet exchange or cache can be operationally important and invisible to a simple AS lookup.
TraceMON combined RIPE Atlas traceroutes with metadata such as autonomous-system mappings, known exchange infrastructure and other hints. The visual interface helped users identify which administrative domains a path appeared to cross and where measurements changed.
Enrichment is a hypothesis process. An IP-to-AS mapping can be stale or ambiguous. An exchange address may be used in a way the dataset does not capture. MPLS can hide hops. Per-flow load balancing can make repeated traceroutes follow different paths. Some routers do not respond, producing gaps.
A good enriched traceroute therefore distinguishes observed data from inferred labels. The hop address and response time are measurement results. The associated AS or facility is an interpretation derived from another dataset. Provenance allows the user to update or reject that interpretation.
The tool reduces the time needed to form an operational question. Instead of staring at addresses, an engineer can ask whether the delay begins after a particular network, whether a path moved through another exchange or whether missing hops correspond to one administrative domain. The answer still requires local telemetry and contact with other operators.
TraceMON illustrates why Candela’s interface work is infrastructure rather than decoration. The design decides how uncertainty is represented and which next steps become obvious. A wrong label can misdirect an incident. A transparent label can accelerate coordination by showing why the system made the association.
RIPE IPmap exposed the importance of maintenance ownership
IP geolocation is often treated as a database lookup. Infrastructure addresses are difficult to place accurately because registration, corporate location and physical router location can differ. RIPE IPmap combined active latency measurements with other signals to estimate where internet infrastructure was located.
Candela worked on the platform and related research, including evaluation of multiple methods. Latency can constrain distance because signals cannot travel faster than physics permits, but routing paths are not straight and queueing adds delay. Hostnames can contain location hints and can be stale. Known infrastructure and operator data can improve estimates and introduce bias.
The methodological value lies in combining evidence rather than declaring one source authoritative. Several weak signals can narrow a location when their assumptions are understood. Ground truth remains difficult: a router interface may serve a link whose endpoints span locations, and an address can move or be reused.
The project also supplies a notable case of maintenance transparency. Candela’s current public profile states that he has not maintained RIPE IPmap since the beginning of 2019 and warns that changes to the active-geolocation platform affected accuracy and coverage. That statement prevents historical authorship from being mistaken for current responsibility.
The boundary matters because services can remain online after their original engineer leaves. Users may cite an old paper while the implementation, probe set and data sources have changed. Current quality has to be evaluated against the current system, not inherited from an earlier result.
RIPE NCC owns and operates its institutional services. Candela’s criticism or disclaimer is evidence about his maintenance role and assessment, not a complete independent audit of the present platform. A responsible article records both facts: he helped design the earlier system, and he no longer controls it.
This episode deepens the profile’s central theme. Interpretation layers need maintenance just as collectors do. A stale enrichment system can produce confident errors. Ownership of data sources, models and code should be visible so users know whose assumptions they are relying on.
BGPalerter changed the interface from investigation to interruption
Candela created BGPalerter in 2019 after leaving the RIPE NCC. The project continuously monitors configured prefixes and autonomous systems using routing and RPKI feeds, then sends notifications when selected conditions are met. His public profile lists the project as current and reports more than 400 installations worldwide; the figure is self-reported rather than independently audited.
The shift from BGPlay is operationally important. BGPlay waits for a user to choose a prefix and examine a period. BGPalerter watches in the background. It can alert on an unexpected origin, loss of visibility, more-specific announcements, unusual paths, RPKI Invalid routes and changes involving ROAs or trust-anchor data.
Continuous monitoring creates a configuration obligation. The system needs an inventory of prefixes, expected origins and allowed changes. A network that acquires a new provider or begins a DDoS-mitigation event may legitimately announce from another AS. If the inventory is stale, the alert is technically correct and operationally unhelpful.
Feed coverage remains bounded. A change can be visible to one collector and absent from another. A collector session failure can resemble withdrawal. The system needs thresholds and source awareness so one missing vantage point does not become a global outage claim.
Notification delivery is another dependency. Email, chat or webhook channels can fail or be throttled. An alerting system should monitor whether alerts were sent and acknowledged. Otherwise, the routing detector can work while the incident process remains blind.
BGPalerter’s open design lets operators run the system themselves and inspect rules. That reduces dependence on a hosted monitoring vendor and transfers responsibility for upgrades, security and feed selection. The project is preconfigured for common use, not zero-configuration. Meaningful deployment requires local knowledge.
The practical objective is not to eliminate the analyst. It is to reduce the time between an observable routing change and a focused investigation. The alert should say which resource changed, which observers saw it and which input produced the conclusion. The responder then checks local routers, change records, reachability and business context.
Routing feeds have to be normalised before a change can become an alert
Public route collectors receive BGP sessions from participating networks. The updates they publish reflect those peer relationships and the collector’s own session state. A monitoring system that consumes several feeds therefore encounters duplicates, delays, resets and differences that are normal properties of the observation system.
The same announcement can arrive from several collectors and at different times. Treating each copy as a separate incident creates noise. Collapsing them too aggressively can erase useful evidence about propagation. BGPalerter needs a model that identifies the resource and event while retaining which vantage points observed it.
Initial state is another challenge. A stream of updates does not necessarily begin with a complete routing table. A monitor needs a baseline against which a withdrawal or origin change can be understood. Collector restarts and peer resets can create bursts that resemble widespread routing events. The system must distinguish loss of the observation session from loss of the monitored prefix.
Timestamps require caution. Collector time, feed transport and processing can introduce delay. The first alert time is not always the first moment the route changed anywhere. It is the first moment the configured monitoring path observed and processed the evidence. Incident reports should preserve that distinction.
More-specific routes complicate grouping. A monitored aggregate may remain visible while a longer prefix appears and attracts part of the traffic. The alarm logic has to decide which prefix lengths are expected and which should trigger attention. Legitimate traffic engineering and mitigation frequently use more specifics, so inventory and context are essential.
AS paths also need normalisation without losing meaning. Prepending repeats an AS to influence selection. Route collectors may show AS sets or confederation-related forms. A path can change while origin remains stable. Whether that matters depends on the operator’s policy and the monitored threat.
The engineering value of Candela’s work lies partly in packaging these details into a system that an operator can run without building a route-analysis platform from scratch. The safety value depends on keeping the details available when an alert is challenged. A notification is useful because it summarises; an investigation succeeds because the summary can be unfolded.
Thresholds translate partial visibility into an operational judgement
Visibility loss is not binary across the internet. A prefix can disappear from one collector, remain present at another and continue serving users. BGPalerter uses configured resources and thresholds to decide when a partial observation should become a notification.
A strict rule can alert on the first missing view. That is sensitive and noisy. A broad threshold can wait until many views disappear and miss a regional problem. The appropriate choice depends on the resource, the feed set and the response cost. Critical anycast services may want regional sensitivity; a small network may prioritise clear global events.
Baselines can be static or learned from recent observation. A static expectation is easy to audit and may become stale. A dynamic baseline adapts and can learn an abnormal state as normal. Change-management integration can improve both by recording planned origin, provider and prefix changes with effective periods.
Thresholds also influence RPKI and path alerts. One Invalid route seen by one collector may indicate a local leak or a global propagation in its early stage. Alerting immediately can be appropriate when the monitored prefix is highly sensitive. Escalating according to additional vantage points can reduce false urgency.
The system’s output should distinguish severity from certainty. A potentially high-impact event can have weak evidence. A low-impact event can be well established. Combining the two into one alarm level hides a useful decision. Responders benefit from knowing both how serious the event could be and how many independent observations support it.
This design work is not visible in a simple project description. It is where monitoring becomes operational policy. Candela supplies defaults and mechanisms, while the deploying network determines which evidence is enough to interrupt a person or trigger another system.
Suppression rules can reduce noise and erase the first evidence of a real event
Continuous monitoring becomes unusable when every planned routing change pages a responder. BGPalerter therefore sits inside an operational process that may include approved origins, expected upstreams, maintenance windows and notification thresholds.
Those controls reduce false positives and create another risk. A broad maintenance window can suppress an unrelated leak. An approved origin can announce a prefix length or path that was never intended. A provider change recorded in a ticket can propagate beyond the authorised scope.
The safer model preserves the event even when the notification is suppressed. Responders can then review what occurred during maintenance and distinguish “not paged” from “not observed.” Rule changes should have an audit trail because they alter what the organisation is willing to notice.
Configuration freshness is part of monitoring health. Prefix inventories, ROAs, providers and contacts change. A tool running current code with stale expectations can generate constant noise or accept a dangerous event as normal.
Candela’s work turns routing observations into usable interfaces. The organisation still owns the policy that decides which observation interrupts a person. That policy needs the same review, expiry and post-change verification as the routing configuration it monitors.
Alert delivery is itself a monitored service
Once an event satisfies a rule, BGPalerter has to reach the people or systems responsible for response. Email, chat integrations and webhooks are convenient and introduce a second availability chain. Credentials expire, channels change, rate limits apply and messages can be filtered.
A production deployment should test delivery independently of real incidents. Synthetic events or health messages can confirm that the route from collector to notification remains open. The system should expose queue and error state so operators can distinguish no alerts from failed delivery.
Deduplication is important. A route can flap and produce repeated transitions. Sending every update can overwhelm responders; suppressing repeats can hide a sustained problem. Grouping events into an incident with a timeline often provides more value than a stream of isolated messages.
Acknowledgement and ownership matter after delivery. A notification in a shared channel does not prove that anyone accepted responsibility. Integration with ticketing or on-call systems can create a record of who is investigating and when escalation should occur.
The alert should carry enough evidence for the first decision: monitored resource, observed origin or path, validation state, vantage points, time and link to further detail. It should not carry so much raw data that the critical change is obscured. Good notification design is another form of visualisation.
Security controls are necessary because alert channels contain network inventory and incident information. Webhooks and tokens can become a route into internal systems. An attacker who can suppress or forge alerts can influence response even without changing BGP.
This operational layer reinforces Candela’s central distinction. Detection is a pipeline, and each stage can fail. Monitoring the monitored network without monitoring the detector creates a new blind spot.
RPKI alerts can distinguish causes only when the changing input is preserved
A route can become RPKI Invalid because the announcement changed, because the relevant ROA changed or because the validator’s view changed. Those causes require different responses. BGPalerter’s value depends on preserving enough provenance to show which input moved.
An unexpected origin with a new Invalid state may indicate a hijack, a customer error or a planned migration whose ROA was not updated. A route that remains unchanged can become Invalid after an address holder narrows a maximum length. A repository or trust-anchor problem can alter validation at scale.
The alert should therefore include time, route, validation source and the relevant authorisation context. A bare “RPKI invalid” message invites responders to treat a classification as an incident conclusion. The classification is a trigger for investigation.
RPKI visibility also differs from service reachability. An Invalid route may still be accepted by many networks. A Valid route can be unreachable for unrelated reasons. Monitoring is strongest when routing observations are combined with active probes and local traffic evidence.
The same restraint applies to origin changes without RPKI. Multihoming, anycast, mergers, provider changes and mitigation services can create legitimate new origins. Approved-change context and maintenance windows can suppress noise without hiding unplanned events.
Alert fatigue is a governance problem. If responders receive repeated legitimate warnings, they learn to ignore the system. Rules should be tuned according to resource criticality and escalation path. A high-confidence unexpected origin may page immediately; a path change may create a lower-priority ticket or enrich another incident.
Candela’s work makes these decisions configurable and visible. It does not assign intent. That boundary protects the tool from becoming an automated accusation system and keeps human validation inside the response chain.
Active probes test reachability that BGP collectors can only imply
A routing collector can show that an announcement is present. It cannot confirm that a user can complete a DNS query, reach a server or avoid excessive delay. RIPE Atlas and similar active-measurement systems fill part of that gap by sending traffic from distributed probes toward a target.
Correlating the two evidence classes is powerful. A prefix withdrawal observed at several collectors followed by failed probes in the same regions supports a stronger outage conclusion than either signal alone. A BGP origin change with stable reachability may still be important and requires a different response. A latency increase without a route change directs the investigation toward congestion, internal routing or the service itself.
The correlation is not automatic. Probe coverage is uneven, and a probe may sit behind local equipment that causes the failure. DNS caching and anycast can send probes to different service instances. A traceroute can change because of load balancing while application performance remains stable. Time alignment and probe selection determine whether the comparison is meaningful.
A monitoring system should therefore treat active results as another partial view. It can choose probes in networks or regions important to the service, maintain a baseline and compare several methods. One failed probe is weak evidence; a coherent pattern across independent probes is stronger.
Candela’s RIPE work supplied interfaces for this kind of reasoning. The value did not come from placing every data source on one screen. It came from helping the user move between route history, latency and path evidence without losing the identity of the measurement.
This layered design is especially useful during a suspected hijack. Public BGP data may reveal an unexpected origin. Active probes can show where traffic is still reaching the legitimate service, where it is failing and where a path has changed. The combined evidence helps prioritise contact and mitigation, while intent remains unresolved.
That method can validate recovery. A route can return before caches, sessions and application paths stabilise. Continued active measurement shows whether service behaviour has followed the control-plane correction. Incident closure should be based on the user-facing result as well as the routing table.
Upstream Visibility compressed several external views into one operational question
Among Candela’s RIPE-era projects was Upstream Visibility, a concise interface for comparing how a prefix appeared from several perspectives. The underlying problem is common: an operator may know its intended providers and still lack a simple view of which upstream relationships public collectors actually show.
A multi-view display can reveal that one upstream is visible only from selected collectors, that a backup path has become dominant or that an unexpected relationship entered the observed path. The interface turns a large set of route records into a question about dependency and reach.
The word upstream is itself context-dependent. The AS path observed from one point can include transit, peering and internal policy choices not evident from the sequence alone. Public data cannot reconstruct every contract. The visualisation supplies routing evidence rather than a definitive commercial map.
This tool sits between BGPlay’s detailed event history and BGPalerter’s continuous notification. It shows Candela experimenting with different levels of abstraction for different tasks. An operator planning resilience may need a summary of upstream diversity. An incident responder may need the exact update timeline. One interface should not be forced to serve both at the same resolution.
The project also illustrates a broader editorial point: a small interface can matter when it removes a repeated analytic burden. Infrastructure value is not proportional to code size. A view that allows an engineer to discover an unintended dependency before a failure can be more consequential than a larger dashboard filled with unrelated metrics.
Open and commercial monitoring systems make different promises
Candela’s projects operate in an ecosystem that includes public data platforms, open-source detectors and commercial observability services. BGPStream and BGPKIT provide programmatic routing-data tools. ARTEMIS combines monitoring with mitigation-oriented workflows. Kentik and other commercial platforms integrate flow, BGP and analytics. ThousandEyes emphasises active internet and application paths. RIPE RIS and RouteViews provide public collector data rather than one incident product.
BGPalerter’s advantage is operator control. A network can run the software, inspect the rules and choose feeds and notification paths. It does not have to send every resource or alert to a hosted vendor. The cost is local operation, upgrades and tuning.
A commercial service can supply broader packaging, support and an integrated data set. It may reduce the work required to correlate feeds and active measurements. It can also create switching costs in query languages, dashboards, historical data and managed response processes.
Public platforms offer transparency and broad research value, but they cannot promise that their vantage points match one operator’s customers. Internal telemetry is more specific and less independently observable. Mature incident detection often combines all three: public views for external evidence, local routers for authoritative internal state and a platform that manages workflow.
The comparison should not be reduced to open versus proprietary. The relevant questions are data coverage, provenance, response time, operational ownership and the ability to verify an alarm. BGPalerter is compelling where a network wants a focused, inspectable detector. It is not a complete replacement for every analytics or mitigation function.
Candela’s career across public infrastructure, open software and a large operator gives him an unusual position in this landscape. The projects show how the same measurement can be packaged for research, public service or production response, with different obligations in each setting.
NTT placed the interface work beside a Tier-1 operating environment
Candela’s current public profile identifies him as a Principal Engineer at NTT, working on the collection, analysis and representation of large network datasets and on automation and monitoring associated with AS2914. This gives his current work a direct production-network context.
AS2914 is the network identifier associated with NTT’s global backbone. The public evidence does not reveal the internal monitoring architecture, team boundaries or operational outcomes. It would be inaccurate to attribute every NTT routing tool or decision to Candela.
The supported significance is narrower. His earlier work on public measurement and visualisation now sits beside the needs of a large operational network. A Tier-1 environment has many peers, customers, routes and changes. False positives consume valuable attention. Delayed detection can affect a broad customer base. Interfaces need to integrate with automation and incident workflows rather than remain research demonstrations.
Production context can improve an open-source project by exposing scale and failure conditions. It can also create private knowledge that does not appear in public code. BGPalerter should not be treated as a complete picture of NTT’s systems, and NTT should not be treated as the owner of every project Candela maintains.
The role also underscores the difference between measurement and control. Monitoring AS2914 can identify a change and supply evidence. Changing routes automatically involves authorisation, safety checks and rollback. Public sources support the monitoring and automation description, not a claim that Candela’s tools autonomously govern the backbone.
His current title is strong evidence of professional standing. It is not a measure of the impact of one project. The value of the profile lies in the continuity between research interfaces, public infrastructure and production operations, with the institutional boundary kept intact at each stage.
Geofeeds let operators publish location and leave trust with consumers
Candela co-authored RFC 9092, which describes how networks can publish geofeed data for IP prefixes. He later created geolocatemuch.com to monitor adoption and configuration. This work addresses a recurring source of operational and commercial error: databases that place addresses according to registration or inference rather than the operator’s intended service location.
A geofeed is an operator assertion. It can provide a mapping between prefixes and geographic information in a standard form. Consumers such as geolocation providers decide whether to retrieve, validate and use it. Publication does not force acceptance.
The mechanism improves transparency because the network can state its own information rather than relying only on third-party inference. It also creates maintenance obligations. Prefixes move, service regions change and broad records can misrepresent diverse users. A stale geofeed can become another source of error.
Monitoring adoption is useful because a standard’s existence does not prove use. A public site can show which networks publish data, whether references are reachable and where formatting problems occur. The evidence remains bounded by what the monitor can discover and by whether downstream databases ingest the feed.
Geofeeds do not solve every geolocation problem. An address can serve users across regions through anycast or distributed systems. The desired location may differ by application: legal jurisdiction, network endpoint or customer market. Operator-published data should be one input among others, with provenance and date.
This work fits Candela’s broader method. It creates an interface where the party closest to the operational fact can publish structured evidence, while consumers retain the decision to trust and combine it. The standard narrows ambiguity without manufacturing universal ground truth.
Vantage-point diversity determines whether an alert describes the internet or one observer
A BGP event is never observed from nowhere. Route collectors receive feeds from specific peers in specific locations and policy contexts. An announcement visible at one collector may be absent at another because the route was filtered, not selected or never propagated to that part of the network. BGPalerter and BGPlay inherit those boundaries from their input.
The number of feeds is therefore less informative than their diversity. Ten sessions from similar networks can provide less independent evidence than a smaller set spread across regions, tiers and routing relationships. A monitoring system should preserve which sources saw the event, when they saw it and whether a source itself became unavailable.
Visibility loss is especially ambiguous. A prefix disappearing from one collector can indicate a withdrawal, a session reset, a collector fault or a policy change between the origin and that observer. Broad loss across independent feeds is stronger evidence of a routing problem and still does not establish the cause.
Origin-change alarms have the same structure. A new origin seen only through one path may be a leak or a collector anomaly. A new origin propagated widely may be a legitimate anycast or provider change. RPKI can add authorisation evidence when a relevant ROA exists, while an Invalid state still needs context such as maximum length, record timing and planned changes.
Candela’s interface work is valuable because it can expose these observations in a form a responder can compare. The danger begins when the interface compresses source diversity into one severity colour without retaining provenance. A concise alarm should be a doorway to the underlying feeds, not a replacement for them.
This has practical consequences for service objectives. A monitoring team should track feed freshness, session resets, collector delay and the share of configured resources seen by each source. The alerting system itself needs an alarm when its evidence narrows. Otherwise silence can be misread as stability.
Vantage-point limits also explain why active measurement complements routing data. RIPE Atlas probes can test reachability or latency from places that do not contribute BGP feeds. Traceroutes can show a changed path without proving the exact interdomain policy. Combining the signals increases confidence only when their different observation models remain visible.
The disciplined conclusion is proportional. One observer establishes that one observer saw a change. Several independent observers establish broader propagation. Local router and traffic evidence determines what the operator should do. Candela’s tools reduce the time between those steps without erasing them.
An anomaly becomes an incident only after local evidence closes the gap
A BGP alert is evidence that selected observers saw a change. It does not say why the change occurred or whether users were harmed. The distinction is central to responsible routing monitoring.
The first response should establish scope. Which collectors saw the event? Was the prefix visible elsewhere? Did local routers receive or export the change? Are active measurements failing? Has traffic shifted? One vantage point can reveal an early signal and cannot support a global conclusion by itself.
The second step is change context. Routing teams should compare the event with maintenance records, provider actions, customer requests, DDoS mitigation and RPKI updates. An unexpected origin may become expected once a planned service is identified. Conversely, a change recorded as planned may have propagated beyond its approved scope.
The third step concerns intent and impact. Malicious intent is rarely visible in the BGP update. A typo, stale configuration and deliberate hijack can produce the same route pattern. Impact depends on which networks accepted the route and whether traffic followed it. Public collectors and active probes can estimate exposure; local telemetry and counterpart contact refine it.
Only then does response begin. The operator may withdraw a route, contact a provider, correct a ROA, adjust filters or communicate with customers. Monitoring software can notify and enrich. Automatic mitigation requires separate controls because a false positive can create the outage the detector was meant to prevent.
Candela’s portfolio is valuable precisely at the transition between the first signal and the focused investigation. BGPlay reconstructs history. TraceMON adds path context. Atlas tools test reachability and delay. BGPalerter pushes the change to the responder. No single tool completes the chain.
This layered view prevents visual certainty from becoming operational overconfidence. The interface should make the next question easier, not make the user forget that another question remains.
Visualisation choices are part of the evidence model
The design of a network interface determines which differences are visible. A graph can emphasise origin changes and downplay update volume. A map can suggest geographic precision that the method does not support. A timeline can make two events look causally connected because they occur close together.
Candela’s work is a useful case for treating interface design as an analytical method. Layout, aggregation, labels and animation are not neutral decoration. They encode assumptions about the object under study.
A responsible tool exposes uncertainty at the same level as the result. Missing collectors, unknown hops, stale mappings and confidence should be available without forcing the user into raw data. The interface can remain clear while showing that the evidence is partial.
Reproducibility is another control. A user should be able to identify the time range, resource, measurement or feed used to generate a view. Shared links to a specific state help incident teams discuss the same evidence. Export or API access lets analysts test an alternative representation.
Visual systems also need accessibility and performance. A graph that works for one prefix may become unreadable during a large event. Progressive disclosure, filtering and stable colour semantics can prevent overload. These are engineering decisions with operational consequence.
The best measure of success is not whether a visualisation looks sophisticated. It is whether an operator reaches a correct, testable next step faster and can explain why. Public user studies and incident case reports would strengthen evidence for that outcome; the current record establishes the tools and their methods more clearly than their quantified effect on response time.
Research training shaped a method that remained useful in operations
Candela’s early work at Roma Tre University and later PhD at the University of Pisa provide more than a chronology of degrees. They help explain why his tools treat visualisation as a testable analytical layer rather than a reporting afterthought. Research requires a method to be described, evaluated and compared. Operations requires the same method to produce an answer quickly enough to matter.
BGPlay emerged from work on dynamic graph representation. The design problem was not simply to draw AS paths. It was to preserve time, reduce clutter and allow a user to inspect transitions at several levels of abstraction. Those are research questions with direct operational value.
His later geolocation work also reflects method discipline. Instead of assuming one database was authoritative, the system combined latency, naming and infrastructure hints and evaluated them against available ground truth. The resulting estimates were conditional on probe placement and data quality. That conditionality is essential in production, where a confident but unexplained location can be worse than an explicit range.
The PhD period overlapped with professional work, connecting academic evaluation with systems already used by operators. Public evidence does not justify assigning every publication or tool to one institution, but it supports a career in which research and engineering reinforced each other.
This background also explains the caution required around adoption metrics. A self-reported installation count is evidence of claimed reach, not a controlled study of effectiveness. A paper result belongs to its data set and method. A visual interface can be useful without proving that it reduces incident time in every network. Candela’s strongest record is the repeated construction of tools and the transparency of their data sources, while quantified outcome claims remain limited.
Open monitoring depends on labour that does not appear in the alarm
BGPalerter is publicly available and has no disclosed standalone revenue or audited project budget. That does not make its maintenance costless. Feed changes, dependency updates, security review, documentation and user support all require time. The more networks rely on the project, the more consequential that hidden labour becomes.
Candela’s employment provides professional continuity, but public sources do not show how his time is divided between NTT work and independent maintenance. Users should not assume that an employer guarantees support for an external project. Nor should they assume that a popular repository has enough reviewers to absorb succession.
A sustainable open detector needs more than occasional feature contributions. It needs people who understand the event model, tests against feed changes, release procedures and a security channel. Documentation should allow an operator to diagnose the detector rather than only configure it.
Institutional services solve the problem differently. RIPE NCC can assign teams and budgets to Atlas, RIS and RIPEstat. Commercial platforms charge customers for support and operations. An independent open project relies on a mix of maintainer time, users and contributors. Each model has strengths and failure modes.
Networks that depend on BGPalerter can improve resilience by contributing reproducible fixes, testing releases and documenting integrations. Funding general maintenance may be more valuable than paying for one private feature. A private fork can solve an immediate need and create a long-term upgrade burden.
The economic value of monitoring is also hard to quantify. Faster detection may reduce outage time, but the saving depends on incident frequency, response and customer impact. Public evidence does not support a universal return figure. Leadership should justify investment through its own risk and operating data rather than assigning a market value to the open project.
This labour question completes the ownership argument. A tool’s source may be public, its feeds may be public and its continued usefulness may still depend on a small number of people. Making that dependency visible is part of responsible observability.
Ownership and maintenance must be stated tool by tool
Candela’s career crosses universities, the RIPE NCC, open-source projects and NTT. The institutions matter because they determine who operates and maintains each system now.
BGPlay and RIPEstat are associated with RIPE NCC services even though the design originated in Candela’s research and engineering. RIPE Atlas, RIS, DNSMON and related platforms are institutional infrastructure. RIPE IPmap continued after his departure, and his public disclaimer draws a clear maintenance boundary.
BGPalerter is his current open-source project, with a wider contributor and user community. NTT’s internal systems belong to the company and its teams. geolocatemuch.com is a separate public project. PacketVis appears on his current site as another product or service association, but the available evidence is insufficient to state its ownership, revenue or customer base.
These distinctions protect both the subject and the reader. Historical contribution should receive credit without making a former engineer responsible for later service quality. A current employer relationship should not be converted into personal ownership. Promotion of a product should not be treated as proof of a financial structure.
Funding is similarly distributed. RIPE NCC services are supported through the organisation. NTT funds its operations. BGPalerter has no published standalone revenue or audited budget. Candela’s compensation and any consulting relationships are not public and should not be estimated.
The maintenance lesson from RIPE IPmap applies across the portfolio: every interpretation layer needs a named owner, current data sources and a change history. An interface can outlive its original designer. Users should know whose assumptions they are running today.
Candela’s durable contribution is a disciplined route from signal to judgement
Massimo Candela’s work does not eliminate the uncertainty of internet measurement. It organises that uncertainty so an operator can act without pretending to know more than the data supports.
BGPlay turns update sequences into a navigable history. RIPE Atlas interfaces make active measurements usable in real time. DNSMON and LatencyMON connect distributed probes with service behaviour. TraceMON enriches incomplete paths. RIPE IPmap demonstrates both the promise of combined evidence and the need to track maintenance ownership. BGPalerter changes routing observations into notifications. Geofeed work gives operators a structured way to publish location claims.
The common mechanism is interpretation with provenance. Each tool reduces complexity and should preserve the route back to the observation. That balance is difficult. Too much detail defeats the interface. Too little creates false certainty.
Candela’s current NTT role suggests that the method remains grounded in production requirements, while public evidence stops short of describing the company’s internal systems. His strongest profile is therefore not that of an inventor who solved BGP monitoring. It is that of an engineer who has spent a career designing the handoff between distributed data and human judgement.
The handoff is infrastructure. During an incident, it determines what the operator sees first, which hypothesis receives attention and how quickly teams move from a public signal to local proof. The software cannot make the decision for them. It can make the decision accountable to evidence.
Member Briefing
Deeper Profile Context
Sign in with the right membership level to unlock the full briefing and source notes.
Only for Strategic Circle
Strategic Circle
Open to all readers. Unlock profile briefings after joining and signing in.
Join Strategic CircleOnly for Leadership Alliance
Leadership Alliance
For qualified IP-asset owners and management; sign in to unlock alliance briefings.
Join Leadership Alliance
