Summary
- Nick Feamster is a professor at the University of Chicago, measurement researcher, institution builder and NetMicroscope co-founder; he makes hidden Internet behaviour usable for accountable decisions.
- His rcc work reported more than 1,000 previously undetected faults across 17 autonomous systems; the Routing Control Platform helped establish a network-wide control model later linked with SDN.
- Projects on spam, broadband, censorship and smart homes simultaneously revealed measurement limits: reputation can misattribute, remote measurement create risks, and encrypted metadata expose behaviour.
- His current ML research asks after efficient, auditable and safe production decisions; his contribution is collaborative systems that preserve uncertainty rather than conceal it.
rcc: checking the combined configuration before deployment
The 2005 paperDetecting BGP Configuration Faults with Static Analysis, published with Hari Balakrishnan, introduced the Router Configuration Checker rcc. It classified persistent faults into two categories. Route-validity faults arise when the control plane selects a route that corresponds to no usable data-plane path. Path-visibility faults arise when a usable path exists but the routers that need it do not learn it. The categories linked configuration commands to consequences understandable to operators.
rcc analysed multiple router configurations and network-wide conditions. The publication reported on 17 autonomous systems, more than 1,000 previously undetected faults and more than 65 downloads by operators. Those are historical figures reported by the authors. They demonstrate neither a universal product nor continued use. They do, however, show real configurations and deployment outside the lab.
Operationally important is the form of evidence. A checker can identify a configuration relationship that violates an invariant before a failure occurs. That differs from a dashboard that reports packet loss after degradation. The engineer receives a reasoned link between a rule and a fault class. This supports review, testing and conversations between teams each with a local view.
The limits are equally instructive. rcc checked only encoded properties and understood syntax. It could not know undocumented business intent, rule out vendor bugs or physical failures, nor turn a clean configuration into proof of the absence of transient disturbances. A network can pass every existing check and violate a requirement never formulated.
This is why rcc stands at the beginning of the profile. It established the ambition to make infrastructure behaviour computable before harm occurs, and the restraint to tie correctness to inputs and specified properties. Modern AI systems face the same test; their limits are harder to see when the model delivers fluent explanations instead of explicit invariants.
The internet usually knows what it did but cannot explain why
A packet reaches its destination or fails. A video call stutters. A domain returns an unexpected address. A mail server refuses a connection. A smart speaker contacts a remote service at a revealing time. Every event leaves traces, but the internet has no central registry with a single authoritative explanation. Its behaviour emerges from independently operated networks, vendor-specific configurations, private interconnection agreements, home devices, application design and varying demand.
This structure prevents any single operator from controlling the whole system. It also complicates diagnosis. A router shows its own routes but not every consequence of the combined routing policy. A speed test measures a transfer without cleanly separating Wi-Fi, access capacity, latency, interconnection and application effects. A censorship probe observes a failed retrieval without identifying the responsible person or institution. A traffic classifier assigns a label without proving it remains correct after a network change.
Feamster’s research can be read as a series of attempts to narrow these gaps. The technique changes, the method remains recognisable: select a relevant hidden behaviour; find an observation point where it leaves a measurable signal; construct a representation that translates the signal into a question for operators, policy or users; finally, test where that representation fails. The last step is crucial. A system that answers confidently without showing its limits can make infrastructure less, rather than more, accountable.
Thus Feamster is a subject of digital infrastructure, although he owns no fibre, operates no public autonomous system and runs no hyperscale cloud. His work lies in the information layer around those assets. It influences route checking, abuse detection, broadband quality description, censorship documentation and the use of statistical models in network operations. Routers and cables move traffic; measurement and analysis decide whether anyone can explain what is happening.
The most compelling narrative is therefore neither a list of prizes nor a claim that one researcher invented multiple fields. It is the story of an observability programme that changed its entity of study while preserving the same discipline: separate direct observation from conclusion, and distinguish experimental result from operational guarantee.
One career, several institutional identities
As of the research reference date of 3 August 2026, the University of Chicago listed Feamster as Neubauer Professor of Computer Science and Faculty Director of Research at the Data Science Institute. His own page also named him Director of the Network Operations and Internet Security Lab, Co-Director of the Internet Innovation Initiative, Co-Leader of netml.io and Co-Director of the AI and Policy Pillar. A DSI post from September 2024 used Director of Technology Policy. These roles can coexist because university functions overlap and change; they must not, however, be merged into a timeless title.
The distinctions are substantive. As a professor he researches and teaches. In NOISE Lab and netml.io he works with students and partners on routing, measurement, privacy and ML. Internet Innovation and Internet Equity provide foundations for public and infrastructure decisions. As co-founder and CEO of NetMicroscope he is involved in a private company aiming to commercialise network quality analytics. His biography also lists expert witness work in technology disputes.
No role automatically transfers the authority of another: a professor is not a regulator, a startup CEO does not turn university research into a customer endorsement, and an expert report is not a court judgment.
This separation fits the content of the research. Feamster’s best work asks which system holds which evidence and what conclusion it supports. The same standard applies to his own description. His current biography credits him with LookSmart’s first web crawler and contribution to Damballa’s first botnet detection algorithm. These are useful, attributable facts about early industry work. They do not prove exact employment dates, sole authorship, contribution shares or complete product history.
The public record on his professional work is considerably richer than on his private life. It reliably establishes neither date nor place of birth, nationality, family background, compensation, NetMicroscope equity, personal assets or wealth. An evidence-bound profile should not speculate from these gaps. The career is substantial enough without adding a celebrity biography unsupported by the sources.
MIT, a web crawler and a dissertation on faults before failure
Feamster remained at the Massachusetts Institute of Technology from undergraduate through doctorate. He received an SB in 2000 and an MEng in 2001 in electrical engineering and computer science, and earned his PhD in 2005 under Hari Balakrishnan in computer science. The titleProactive Techniques for Correct and Predictable Internet Routingplainly sets out the early programme: rather than wait for a routing fault and its outage, the network should expose enough structure to verify important properties before deployment.
The early work at LookSmart is a useful, bounded precursor. A web crawler must explore a large graph that changes during observation. It encounters dead links, duplicates, conflicting responses and unreachable areas. The crawler did not lead directly to the later routing systems; there is no basis for that. The continuity is methodological: a distributed system is not visible from a single local view. Useful knowledge requires systematic collection and an explicit representation of what was found.
Damballa offered the adversarial variant of the same problem. Botnets conceal membership and control. Feamster’s official biography says he contributed to the company’s first botnet detection algorithm. Public evidence does not reconstruct the full team or later product use. It does show, however, an early shift between academic systems research and an operational security firm that had to infer malicious coordination from network traces.
The dissertation emerged in a world where interdomain policies were distributed across devices and organisations. Router commands described peering, customer relationships, export policies, backup paths and traffic engineering. Every configuration could be individually plausible while the overall network hid a loop, a black hole or an unwanted route. The problem was not just a flawed protocol but the difficulty of reasoning about a distributed programme built from many local rules.
This framework remained formative. Feamster often treated the operational system around an algorithm as the real research unit: configuration files, measurement points, data pipelines, interfaces, operator incentives and institutions. That extended the work beyond a single theorem or classifier. It also enlarged the responsibility: once a system touches real operators, households or people under censorship, deployment and ethics become part of technical quality.
BGP configuration as a distributed programme
The Border Gateway Protocol lets autonomous systems exchange reachability while preserving their business and routing policies. That autonomy joins networks with different owners and goals. It also means that no single engineer designs the global outcome. Policies are set indirectly through announcements, preferences, filters and internal route distribution.
Even within a single autonomous system the task is large. Hundreds of routers may distribute external routes internally along multiple paths. A route learnt at an edge must be visible where it is needed, but not everywhere. Export rules should protect customer or peer routes from the wrong neighbours. Backup paths must appear after failure without loops or persistent oscillation. The operator knows the business intent; the devices hold the executable fragments.
Feamster’s early work treated those fragments as an analysable programme. Router configuration went from text to the input of a network-wide computation. Correctness could be expressed as an invariant: a chosen route should lead to a usable forwarding path; a usable path must be visible to the routers that need it; export rules must preserve the intended relationships; internal distribution must not create persistent inconsistency.
The software-analysis analogy shifted the point of intervention. Classic troubleshooting begins after a symptom. Static analysis asks whether a known fault class is already present in the configuration. It needs neither test traffic nor a customer complaint. For critical networks, moving a defect from the incident queue to the review queue can be more valuable than speeding up later diagnosis.
The analogy has limits. The network state includes live routes, topology, vendor behaviour, transient convergence, hardware tables, broken links and undocumented business knowledge. A static checker can be exactly right about its model and miss a fault that lies outside it. Later research returned repeatedly to this difference between a useful representation and the complete operational world.
The Routing Control Platform moved decisions out of individual routers
rcc asked whether distributed configuration satisfies known conditions. The Routing Control Platform asked a different question: why should each router reconstruct route-selection information on its own? Classic iBGP distributed external routes via a full mesh or route reflectors. The mesh became unwieldy as networks grew. Reflection scaled better but could hide routes, create unexpected decisions and complicate analysis.
RCP proposed a logically centralised service that collects external BGP routes and internal topology, selects routes for individual routers and communicates them via ordinary iBGP. The forwarding devices remained unchanged. What changed was the location of the choice and the view available.
“Logically centralised” did not mean a single fragile physical box. The service could be replicated and distributed yet still present a coherent decision function. The same holds for later SDN. A controller can act network-wide without living on a single machine. The engineering questions become state consistency, recovery and safe interfaces, rather than a central-versus-distributed dichotomy.
RCP respected installed infrastructure. It required neither a new forwarding plane nor immediate hardware replacement. That matters in networks with long-running equipment, contracts and procedures. Research gains practical value when it reaches operations through a known interface.
The evaluation used real backbone data but does not support a broad production count. It is fair to say: RCP demonstrated a functional and influential architecture. It is not fair to say: RCP replaced internal BGP industry-wide. The NSDI Test of Time Award 2015 attests to intellectual endurance, neither market adoption nor ownership of later controller ideas.
A significant SDN contribution, not a sole-inventor story
SDN is often told as a clean break: control moved into software, forwarding became programmable, a new era began. The actual history is messier. Active Networks, network virtualisation, 4D, Ethane, RCP, OpenFlow, NOX and other projects solved different pieces of programmability, control separation and network-wide management. Feamster was an important contributor but not the sole inventor.
RCP showed that a control service with a broader view could compute route decisions and deliver them to existing equipment. rcc showed that network policy could be checked against invariants. Together they moved the question from “Which command sits on this router?” to “What behaviour does the whole control system implement?” That shift is an intellectual foundation of programmable networks.
Feamster later co-authoredThe Road to SDN, which described the field as a collection of ideas rather than a moment of invention. This view resists founder myths. A field emerges when research groups, operators, vendors and standards communities solve neighbouring problems and make them deployable.
Why (and How) Networks Should Run Themselvesfrom 2017, with Jennifer Rexford, extended the argument to continuous operation. High-level intent drives decisions, telemetry shows results, the system adjusts. The provocative title does not prove that engineers are superfluous. Self-adjustment needs correct objectives, trustworthy telemetry, safe action, limited privileges and a stop when evidence is unclear.
Today this agenda is an AI operations problem. Language models propose configurations, summarise incidents and select tools; they can also invent causes, misunderstand policy and exercise too much authority. The early routing work sets a hard standard: learned recommendations need explicit checks and observable consequences, not just convincing style.
Spam made the network around the message more important than the message
At Georgia Tech the observation target shifted from configuration faults to adversaries. Spam campaigns and botnets were agile. Compromised machines came and went, addresses changed, domains were replaced, control was distributed. Content signatures spotted known messages; the delivery infrastructure often showed the more durable pattern.
The SIGCOMM paperUnderstanding the Network-Level Behavior of Spammersfrom 2006 analysed, according to the publication, more than ten million unwanted messages. It examined origin, activity duration and relationships to address space and routing. What mattered was not an eternal percentage but the view of abuse as infrastructure: population, routes, time and concentration revealed coordination that text alone did not show.
Network features are available early in a connection, before full acceptance or inspection. That can save processing, enable scale and partly preserve payload privacy. At the same time, the abstraction carries risks. A residential prefix contains innocent users and compromised devices; shared hosting serves legitimate and malicious domains; addresses change. Infrastructure reputation is only usable with uncertainty, ageing and recourse.
DNSBL Counter-Intelligence exploited the adversary’s defensive behaviour. Botmasters checked DNS blocklists for their machines. Typical queries could hint at bot membership. Adversarial reconnaissance became a signal but remained heuristic: a query could be benign and a suspicion list needs confirmation.
SNARE rated senders using spatio-temporal and network features during SMTP. The evaluation reported around 93 per cent accuracy at a low false-positive rate. The figure belongs to a dataset, threat and threshold and does not stay constant for 17 years. Adversaries, mail infrastructure and distributions change. What was proved was the utility of early signals, not a permanent benchmark.
Dynamic DNS Reputation extended the logic to domains. Registration, name servers, address changes and resolution can distinguish harmful infrastructure from stable services. The classifier assigns risk before knowing any payload. This prefigured later network ML: represent metadata, learn a decision rule, handle the consequences of incomplete representation.
Reputation is an operational decision, not an objective label
Behavioural clustering of HTTP malware went further. Instead of an exact binary signature, it grouped samples by communication behaviour and generated network signatures. This helps when code changes but protocol, destination or timing stays the same, but can also lump unrelated traffic together if the representation is too coarse.
The difference between signal and judgement determines use. A model reports similarity to known abuse; the operator decides on observation, rate limiting or blocking. Blocking an entire prefix can hit innocents. Technical design includes threshold, age of evidence, scope of action and correction path.
Early Damballa work and a patent assigned to Georgia Tech Research Corporation show commercial and IP context. The patent names Feamster, David Dagon, Wenke Lee and others as inventors for detecting and responding to attacking networks. It is formal evidence of inventorship and assignment, not proof of sole creation, production, licence revenue or any legal validity.
The larger contribution was treating reputation as an infrastructure problem rather than an accuracy metric. A defender needs a classifier that runs in time, at available line rate, on lawfully collected data, with manageable errors. A study optimises one piece; a production system carries the entire chain for years against a shifting adversary.
Thus the security work leads to today’s ML programme with a stronger focus on feature cost, drift, privacy and deployment. The question remains: what evidence makes a decision inferred from network signals reliable enough to affect real traffic?
Broadband measurement at the gateway changed the observation point
Broadband complaints are easy to voice and hard to diagnose. “The internet is slow” can denote a thin access link, poor Wi-Fi, a busy home device, a congested interconnection, a distant server, application delay or an unsuitable test. Measurements from a laptop inherit its software and local network. Measurements in the provider core miss what the household experiences.
The gateway approach placed controlled measurement at the boundary between home and ISP. The 2011 SIGCOMM study used long-term data from nearly 4,000 gateways across eight providers within a larger stock of over 4,200 devices. It examined throughput, latency, access technologies and traffic shaping. Analytically valuable was the view of the access service with less endpoint variance.
The answer still did not become automatic. The gateway shares an environment with devices and Wi‑Fi. The measurement server has its own path and capacity. Tests can disturb household traffic or run at untypical times. The method improved attribution, not a complete map of user experience.
BISmark turned this into a reusable testbed. The 2014 USENIX ATC paper described customized routers and a backend for measurements and applications. At the time, the system ran in hundreds of households across roughly 30 countries and was used at nine institutions. Those time-bound numbers show infrastructure beyond a single dataset.
Running a home testbed is its own infrastructure work: ship and support hardware, handle powered-off devices, ageing firmware, drifting clocks, intelligible consent, stable data schemas and pipelines. The published system therefore demonstrates institution building as much as measurement design.
The programme also shows the policy significance of method. A controlled gateway result is not the same as a browser test, an ISP counter or an advertised tariff. Each measures a different segment. Public decisions are more robust when measurement position and limits are visible, rather than disappearing behind a single speed number.
With faster access, speed alone no longer explained the experience
Early broadband policy often asked whether the advertised rate was delivered. That remains important when capacity is scarce. It explains less once latency, Wi‑Fi, content delivery and application design dominate the experience.
A study of more than 5,000 broadband networks found that above a study-specific range raw access throughput was often no longer the sole bottleneck. A faster tariff does not necessarily make a page faster when round-trip time, entity dependencies or server behaviour determine completion. The threshold is not universal; what endures is the multidimensionality of quality.
Reliability is another dimension. A service can have a high median rate and disappoint households through brief outages. A video call, exam or telehealth session can be disrupted by a short interruption that disappears in monthly averages. Long-term measurement becomes necessary.
Encrypted DNS research showed similar trade-offs. Resolver and protocol choice affect latency, privacy and reachability. In the measurement panel no single configuration was optimal for all users and networks. A security improvement can cost performance in one setting and not in another. A benchmark is not a universal prescription.
COVID measurements showed rapid operational change. Participating ISPs experienced abrupt shifts in traffic and interconnection, capacity expansion and altered usage. The results applied only to the networks measured but demonstrated the weakness of planning that only knew stable historical averages.
Here broadband measurement becomes quality-of-experience engineering. Operators must connect throughput, latency, loss, outages and interconnection to video stalls or conference quality. This line entered the product thesis of NetMicroscope. Academic results do not prove every company claim, yet the technical lineage is direct.
Interconnection measurement turned a business dispute into a shared data problem
The path between a household and an application crosses multiple corporate boundaries. An access provider exchanges traffic directly, at an Internet exchange or via transit. Congestion can occur on a single link, in one direction and at a given time. Public debates about “interconnection” therefore conflate different conditions.
The Interconnection Measurement Project had participating ISPs install a common tool on interconnection links. It reported roughly 2,900 links and about 31 per cent peak utilisation for June 2021, as well as capacity augmentations. Aggregates do not prove that every user path was free. Averages hide brief spikes, individual routes or local failures. Valuable was the shared method and language.
A shared system lessens one dispute and creates another. Entities must agree on links, sampling, privacy and publishable summaries. Providers have commercial reasons for restraint. Researchers need enough detail without exposing customer relationships or security-sensitive topology.
Feamster’s role was building measurement capability and institutional collaboration, not regulation or compulsory participation. The project led from the home gateway to an evidence system shared with operators. The technology was inseparable from governance: useful data required the cooperation of link owners.
Internet Equity extended performance with access, affordability and reliability
A broadband line can be present and unaffordable. A household can subscribe and receive unstable service. The provider meets the rate while Wi‑Fi, building wiring or the device impairs the application. A single coverage map hides these differences.
The University of Chicago’s Internet Equity Initiative joins availability, infrastructure, affordability, adoption, performance and reliability. Portals combine network measurements with demographic and policy data. Devices in Chicago supplied direct performance evidence; public datasets enabled comparisons.
The method shows patterns, not every cause. Differences may relate to income, building type, competition, tariff, equipment or historical investment. Demographics support investigation but do not prove causality. As with routing, a representation makes the question testable but does not replace missing variables.
In 2025 UChicago reported that Internet Innovation work contributed to Illinois planning for 175,000 underserved households, businesses and community anchors. That is a material institutional statement involving the Illinois Broadband Lab, the state and partners. It must not be rewritten as a professor’s personal connection achievement nor as proof of completed construction.
From gateway to equity the measurement user changed: operators diagnose lines, cities target support, states use validated data in grant programmes. With policy weight comes higher demands for transparent definitions, versioned data and the presentation of remaining uncertainty.
Censorship research turned observability into a matter of human safety
Censorship is hard to measure for the same structural reason that routing is hard to explain: what is observed is a result of multiple systems. A failed DNS resolution can be a state filter, a local firewall, an ordinary error, a server outage or routing instability. A reset can be injected, generated by an endpoint or caused by an apolitical middlebox. The signal rarely comes with a signed attribution.
The consequences differ because measurement can expose people. Researchers need observation points inside censored networks; volunteers may face reprisal. Remote methods reduce local recruitment but can involve users, websites or systems that never consented. The safety model is part of the measurement architecture.
Feamster’s work began in 2002 with Infranet. The system treated censorship as a circumvention problem. Cooperating web servers encoded covert requests inside normal-looking HTTP and hid responses in images. The web assumptions are historical; the idea remained: censorship is a contest over which traffic patterns can be told apart from ordinary communication.
Later the work shifted from circumvention to measuring the block. Large, repeatable measurements could document undisclosed filters. At the same time a difficult ethical boundary arose: a public evidence system can shift risk onto the single machine that generates the signal.
This tension is central. Making the internet explainable can harm people who never asked the question. Ethical quality depends on target selection, consent, rate limits, data retention, disclosure and plausible reprisal, not just on statistics.
Encore showed how scale can overtake consent
Encore used cross-origin browser requests to test web resources from different networks. A participating website had a visitor’s browser issue a request and allowed a blocking inference from the result. This promised reach without specialist software in every country.
The same mechanism created a serious ethics dispute. A visitor to an unrelated page could unknowingly become a measurement point. A request to a sensitive domain could attract a censor’s attention. A third-party website could probe material it had not chosen. Risk bearer and data recipient were not the same.
No Encore for Encore?criticised consent, transparency, user safety and possible harms to third-party sites. The paper also documented the dialogue with Feamster and changes to the system. It is neither a finding of research misconduct nor a footnote erased by later work. It identified a real design risk in a legitimate public project.
Feamster and Ben Jones subsequently publishedCan Censorship Measurements Be Safe(r)?. The title itself treats safety as a continuum. Coverage, repeatability and accuracy stand against exposure of volunteers, targets and bystanders. More reach is not better if individual risk is not bounded.
The episode changed the research content. Ethics was no longer an external review but part of the threat model: people, testing organisations and observing authorities belonged to the architecture. That remains relevant for browsers, home devices and AI agents as distributed sensors.
Augur, Iris and later systems expanded reach without hiding risk
Augur inferred connectivity between remote locations from TCP/IP side channels without possessing a classic measurement point at either end. The paper reported validation in nearly 180 countries over 17 days and guardrails against involving individual users. Reach grew, but the inference depended on operating system, address choice, filter asymmetry and statistics.
Iris concentrated on DNS manipulation. Repeated resolver queries were compared across locations. DNS provides structured evidence, but legitimate systems cache, redirect, localise and filter. A deviation is the beginning of attribution, not proof of a particular authority.
Test lists also introduce bias. Checking only global political websites misses local languages and topics. In 2018 a project using NLP and search identified 1,125 websites missing from the then-largest China block list. That improved coverage at the time but remains a time-bound artefact.
GFWeb examined HTTP and HTTPS filtering by the Great Firewall over 20 months for USENIX Security 2024. The paper reported 1.02 billion tested domains and hundreds of thousands of pay-level domains with various filters. Those are measurement counts, not a headcount. They show that protocol-specific tests reveal different system parts and a single method undercounts.
Turkmenistan research reported 15.5 million tested domains, 122,000 censored domains and derived overblocking rules affecting millions more. Circumvention was also studied. Published circumvention helps users and teaches the censor. Disclosure timing and local knowledge are technical decisions with human consequences.
Feamster’s role is significant and collaborative. He co-authored systems, helped build communities and teaches Internet Censorship and Online Speech. He is neither the sole founder of every observatory nor the originator of Geneva through thematic proximity. Attributing papers and roles separately is more precise.
Measurement ethics became part of technical correctness
The censorship work illustrates a general point. A statistically powerful system can fail technically if it cannot be operated responsibly. Consent, target selection, frequency, data minimisation and publication determine whether repetition is possible without unacceptable harm.
Not every risk must be eliminated. Complete safety may be impossible against an adversarial state. What is required are explicit, assigned risks weighed against public benefit. Those most exposed must not disappear behind a global domain count.
The same applies outside censorship. Broadband probes reveal household patterns, IoT tools collect metadata, reputation scores deny service, synthetic datasets can memorise protected traces. Every measurement creates a new infrastructure with users, rights and errors.
Feamster’s career is valuable precisely because of a documented dispute. Encore shows criticism, change and later safety work; later safeguards must not rewrite the original risk. A serious profile can honour learning and preserve the dissent that made it necessary.
Smart homes showed what encryption does not hide
Encryption protects payload content, but a network requires timings, packet sizes, directions and destinations. A smart plug, camera, television or voice assistant contacts predictable services in recognisable patterns. An observer can infer, without message content, that a device was switched on, video was streamed or an event was reported.
A Smart Home Is No Castledemonstrated this side channel;Spying on the Smart Homeadditionally examined traffic-shaping defences. The second paper reported for its test scenario that a constant-rate scheme could hide activity with roughly 40 kilobytes per second overhead. That is not a universal price for privacy. Device mix, threat model, link capacity and desired camouflage change the trade-off.
The work corrected a common simplification: “encrypted in transit” can be true while metadata exposes behaviour. A privacy notice that considers only content omits a relevant risk. Padding and shaping cost bandwidth, energy or latency; the cost may fall on the household rather than the manufacturer.
Practically the question is not whether lab analysis is possible but who can observe the home, which inference is robust and who can change the design. ISP, local attacker, device manufacturer and cloud provider see different things. Measurement shows the leakage; consumer protection demands decisions about defaults, disclosure, remedy and responsibility.
IoT Inspector became a consumer tool and research infrastructure
IoT Inspector took the smart-home work from controlled experiments to an open-source tool for one’s own network. Users selected devices, saw contacted destinations and could, with consent, contribute labelled metadata. The 2020 paper documented thousands of users and tens of thousands of devices from many vendors and categories.
Later reports mentioned 44,956, 54,094, more than 55,000 or about 63,000 devices because time windows and counting methods varied. The largest number without a date would be false precision. What matters is the order of magnitude at which support, label quality, privacy and maintenance became central research tasks.
A user-facing tool also exposes ambiguity. A domain may serve multiple cloud customers. A device label may be wrong. A connection to a tracker does not automatically say what data were transferred or what harm occurred. Visibility is not the same as practical remedy.
The team retrospective addressed incentives, consent, data minimisation and operations. An open research offering can develop duties of a service provider: it stores sensitive hints, depends on entity understanding and must continually communicate limits.
Related studies on selected medical IoT devices, connected toys and user perception extended the work toward consumer protection. Results must stay tied to product, version and date. Firmware changes, manufacturer fixes and different deployments alter the outcome.
Network Machine Learning is a pipeline, not an isolated classifier
Machine learning entered Feamster’s work through spam and reputation long before generative AI. netml.io made the environment explicit: representation, label acquisition, feature extraction, runtime, drift and subsequent action.
nPrint represented packets bitwise in a standardised way; nPrintML connected this to automated modelling. The goal was not the universally best representation but more reproducible comparisons with fewer hidden feature-engineering changes.
Traffic Refinery examined the cost of features at high rate. A statistically accurate model is operationally poor if extraction drops packets, exhausts CPU or answers after the decision window. LEAF handled concept drift from new applications, devices and networks. Production requires criteria for retraining, rollback and environment-specific error monitoring.
CATO jointly optimised prediction and system. The NSDI 2025 evaluation reported up to 3,600-fold lower inference latency and 3.7-fold higher loss-free throughput under its experimental conditions. Those are not general production promises. Important is the joint optimisation of statistical quality and packet-processing cost.
Current work extends to low-cost classification, queueing, probe selection, L4S field measurement and language models for misconfigurations. The 2005 dissertation used explicit invariants; a 2026 model can infer a likely problem from examples and text. It covers cases that are hard to formalise but can replace proof with plausibility if the recommendation is not checked against the observable network.
Synthetic traffic aims to share useful data without revealing the real network
Real traffic traces are hard to share. They disclose communications, users, devices, organisational structure and proprietary applications. Labels are expensive, traces age quickly. Synthetic data seek to generate useful properties without releasing the originals.
NetDiffusion used diffusion models with protocol conditions for packets. NetSSM added state and multiple flows. GATEAU under Feamster and Francesco Bronzino frames privacy, collection cost and label scarcity as a joint problem. The goal is a middle ground between technically valid but unrealistic random packets and real traces that cannot be safely shared.
Generated traces can preserve marginal statistics while losing needed correlations, memorise sensitive examples or satisfy syntax without capturing congestion, session state or user behaviour. A classifier trained on them can perform well in evaluation and poorly in operation.
The 2026 work on privacy–quality trade-offs measures these risks instead of equating “synthetic” with anonymous. Privacy must be tested against plausible attacks, and utility against the intended task. Synthetic traffic is a research instrument, not a certificate that the original network has vanished.
NetMicroscope tests whether the measurement programme can become a business
NetMicroscope is the clearest commercial translation of Feamster’s broadband and ML research. The company lists Feamster as CEO and co-founder, and Francesco Bronzino as CTO and co-founder. University commercialisation material dates the founding to 2021 and describes a remote team around Chicago and Lyon.
The product thesis joins throughput, latency, loss, device state and application signals into an estimate of user quality. An operator may know that a line is active but not why a video degraded. NetMicroscope says ML infers application experience and spots problems before a complaint.
Verified financial data are limited. In January 2024 the George Shultz Innovation Fund awarded $200,000 for product, market, team and intellectual property. The company also participated in I-Corps and Compass. These are early signals, not proof of valuation, total funding, revenue, customers, retention, market share or profit.
The boundary between university and firm deserves ordinary scrutiny, not insinuation. Traffic classification papers may overlap with a product. Readers need disclosure of affiliations, funding, data access, licences and IP. The material shows no exclusive transfer of all BISmark, nPrint or university code.
The conclusion remains modest: NetMicroscope is testing whether years of measurement research can become a paid operational service. What is shown are founders, direction and a university award, not commercial dominance.
Teaching turned the research arc into an infrastructure curriculum
Feamster’s teaching follows the research. At Georgia Tech it covered internet architecture, security, next-generation networking and later SDN. At Princeton networks connected with information security and technology policy. At Chicago it includes Machine Learning for Computer Systems, Internet Censorship and Online Speech, and Security, Privacy, and Consumer Protection.
In 2020 he co-authored the sixth edition of Andrew Tanenbaum’sComputer Networks. His teaching page also credits him with developing and initially running the Computer Networking course in Georgia Tech’s online Master of Science in Computer Science. The online course reached far more than a single campus cohort; the founding claim should nonetheless stay tied to the documented teaching page.
The 2026 Quantrell Award is independent institutional evidence of teaching’s standing. UChicago highlighted open problems, collaborative thinking and realistic exercises. Student testimony is qualitative, but the award shows the career cannot be reduced to papers and startups.
Institution-building widened the audience. Feamster directed Princeton’s Center for Information Technology Policy and later helped lead UChicago programmes joining network metrology, policy and data science. Workshops on free and open communication systems connected ethics and measurement design.
A fair portrayal keeps collaborators visible. Many systems were implemented or led by students and junior researchers. IoT Inspector, censorship projects, broadband testbeds and ML systems have large author teams. A professor can provide direction and institutions without being the sole author of every result.
Policy work supplies evidence but exercises no regulatory authority
Feamster’s policy role is rooted in measurement. UChicago names collaboration with the Federal Communications Commission and the City of Chicago. Internet Equity can evidence access, affordability, reliability and performance. It decides neither subsidies, prices nor network orders.
The boundary matters because technical evidence gains weight inside agencies. A speed test informs a proceeding whose legal criteria are set elsewhere. An interconnection graph explains utilisation, not contracts or routing decisions. A censorship dataset documents interference without completing the legal and political judgement.
Current AI privacy work carries the same problem into a new ecosystem. A Google Privacy Faculty Award supports research on third-party integrations in language-model systems. Users see one interface while prompts, context and inferred attributes travel between plugins, APIs and remote services. As in the smart home, a visible product coordinates hidden data relationships.
The 2026 publication list includes work on implicit LLM inference, where sensitive attributes are inferred without a classical identifier. Data minimisation becomes harder: removing names or account numbers does not prevent the deduction of health, political or demographic information from ordinary interaction.
The policy value lies in visible evidence boundaries. A model, dataset or professor can inform decisions without owning them. Good use demands transparent methods, error estimation, versioned data and a clear separation between measurement result and authorised action.
What the career made visible and what remains hidden
Across routing, spam, broadband, censorship, smart homes and ML, Feamster’s projects transformed diffuse operational problems into evidence systems. rcc connected configurations to invariants. RCP gave route selection a broader view. Reputation systems inferred coordination from metadata. Gateways isolated broadband factors. Censorship systems compared remote signals. IoT Inspector linked device traffic with labels. NetML joined representation, operational cost and drift.
The internet did not become fully legible. A clean configuration does not prevent a physical fault. Reputation does not prove guilt. Speed does not explain affordability. A DNS anomaly does not identify an authority. Encrypted metadata does not reveal every action. Synthetic traces do not guarantee privacy. A language model is not correct because its diagnosis sounds coherent.
These limits argue not against measurement but for a careful operational system around it. Useful evidence shows source, age, scope and uncertainty. Consequential actions need review, recourse or reversal. Researchers should distinguish paper, institutional and replication results. Commercial claims do not inherit academic authority without their own evidence.
Feamster’s enduring contribution is the repeated construction of that middle layer between opaque infrastructure and consequential decisions. It helps operators, researchers, users and institutions ask better questions of systems that should not explain themselves. The result is not certainty but a more disciplined separation of observation, inference and judgement.

