Summary

  • OpenINTEL is a collaborative active-DNS measurement platform operated by the University of Twente, SIDN, NLnet Labs and SURF rather than a registry, resolver or passive-DNS company.
  • Its homepage reports about 308 million domains measured daily, 5.9 billion data points produced each day and 13.6 trillion observations accumulated since regular operation began in 2015.
  • Longitudinal consistency creates the platform’s value, while source selection, vantage, query method, software version, method changes and missing-data controls bound every result.
  • OpenINTEL does not observe user demand or every DNS state; some datasets are open under a non-commercial licence, while zone-access contracts keep other material controlled.

A University of Twente system became shared national research infrastructure

OpenINTEL’s implementation began in 2014 at the University of Twente. The first complete daily run demonstrated that the pipeline could query, process and store a very large namespace within the operational window needed for repetition. Regular measurements began in March 2015. The shift from an experiment to a daily system created the archive’s main value: continuity.

The project was founded by researchers including Anna Sperotto, Mattijs Jonker and Roland van Rijswijk-Deij, whose current project roles span research leadership, data architecture, measurement design and funding. The operating model expanded beyond one university. SIDN brought registry expertise and sustained support. NLnet Labs joined as a DNS-software and research partner. SURF supplied research-network and infrastructure context. The four institutions now jointly operate the project.

This arrangement should not be described as a standalone company. There is no verified OpenINTEL corporation with shareholders, consolidated revenue or a valuation. Staff, hardware, contracts, grants and data rights sit with the partner institutions. The project’s governance is less formal in public than a foundation board, but its institutional diversity reduces dependence on one laboratory.

Each partner also contributes a different view of the DNS. A university values reproducible research and student work. A registry understands zone data, operator relationships and the constraints of access agreements. A DNS-software organisation brings protocol and implementation knowledge. A national research network can support the compute and connectivity required for sustained measurement.

The model creates boundaries. Partners can fund equipment and staff without publishing one project budget. Zone data obtained under contract may be measured but not redistributed freely. Contributor pages can become stale when people change jobs, so they are reliable for project history and less reliable for unrelated current titles. Shared operation does not make every institutional asset jointly owned.

OpenINTEL’s development from laboratory system to common infrastructure also changed its service obligations. Researchers depend on data continuity. Operators need identifiable traffic and a way to report harm. Data users need stable formats and access terms. Storage migrations must preserve history. The project has to behave like a long-lived observatory rather than a paper’s temporary dataset.

The first complete day was a technical milestone. The decision to keep measuring for a decade was the institutional achievement.

The DNS answers the present and forgets the past

A DNS query asks for the current answer available through a particular path. The response may identify name servers, addresses, mail systems, certificate-related records or other configuration. Tomorrow, the operator can change it. The previous state may remain in caches for a while, appear in logs held by one provider or disappear from public view entirely.

That behaviour is appropriate for a live naming system. The DNS does not exist to provide historians with a complete ledger. It exists to map names and other identifiers under distributed authority. Registries, registrars, authoritative operators, recursive resolvers and applications each retain different evidence. No institution naturally preserves a longitudinal view across many namespaces and record types.

The absence matters. Researchers want to know how DNSSEC adoption changed, when hosting or mail infrastructure moved, how concentrated authoritative service became and whether a policy or vulnerability altered behaviour. Security teams want to reconstruct what a domain resolved to before an incident. Policymakers want evidence of dependence on providers. A one-time scan can describe a state; it cannot reveal the transition.

OpenINTEL was built to create that temporal evidence through repeated active measurement. The project obtains or constructs lists of names and address ranges, sends defined DNS queries on a schedule and stores responses with timestamps and metadata. Repeating the process allows a researcher to compare like with like over days and years.

The word “like” requires discipline. Source lists change. New record types are added. Software and infrastructure are upgraded. Some days are incomplete. An authoritative server may rate-limit or block the project. A response can vary by vantage, anycast location or time. Longitudinal value depends on recording those changes rather than treating the archive as one perfectly uniform table.

OpenINTEL’s central contribution is therefore not simply scan volume. Internet-wide scanners can also produce enormous datasets. The distinctive asset is a long-running instrument whose methods, partners and data products are stable enough for change itself to become a subject of study.

That strength supports the title “the DNS’s daily historical record” only with a boundary. OpenINTEL preserves a history of the measurements it was configured to make. It does not record every DNS query, every domain or every answer seen by users. The archive is large precisely enough that careless language can make it sound universal. Its credibility depends on resisting that temptation.

The target list sets the archive’s field of view

Active DNS measurement requires a target population. OpenINTEL can receive zone-derived domain lists, certificate-transparency names, popularity lists, country-code apex data, address ranges and other sources. Each source answers a different research question and contains a different bias.

A registry zone can provide broad coverage of names delegated under that top-level domain, subject to contract. It may omit names in other namespaces and says nothing about whether a domain hosts an active service. Certificate Transparency logs reveal names associated with publicly logged certificates, favouring TLS-enabled services and exposing subdomains not present in zone lists. Popularity lists emphasise frequently accessed names under opaque or changing methodologies. Reverse-DNS measurement begins from address space rather than names.

Combining sources increases reach and the risk of double counting or changing composition. A homepage figure of 308 million domains measured daily should be read as the project’s current metric for configured domain observations, not 308 million unique active websites. A domain can be parked, delegated without content, duplicated across lists or used for mail and infrastructure rather than a website.

Source selection affects longitudinal interpretation. Suppose a new top-level domain is added to the measurement. The total number of observed records rises because the instrument expanded, not because the DNS changed organically. A popularity list can revise its methodology and create apparent churn. Certificate Transparency coverage can increase as issuance practices change. Analysts need versioned lists and inclusion criteria.

The project’s methods can still produce robust trends when comparisons are bounded. Researchers can examine a stable subset of names across time, control for additions and classify source types. The archive’s size allows rare events and infrastructure relationships to be studied, but large numbers do not compensate for an undefined population.

List governance is also commercial and political. Registries may permit measurement under agreements that restrict redistribution. Operators can object to load. A project committed to open science cannot simply publish data it has no right to share. The archive therefore has open and controlled layers.

The first question for any OpenINTEL result should therefore be: which names or addresses were eligible to be measured on those dates? The answer is not background detail. It defines the claim.

OpenINTEL measures names and address space from a Dutch institutional base with global source lists and collaborations. That gives it worldwide subject matter, not automatic representation of every region or user experience.

Zone access varies among registries. Some country-code lists are comprehensive, others are assembled from public sources and some cannot be redistributed. Certificate-derived names favour services using public certificates. A measurement vantage can receive an anycast answer different from one seen in another continent. Split-horizon and geolocated DNS can make both observations valid.

Researchers comparing countries therefore need to separate the domain’s registration label from the location of its operator, users and infrastructure. A .br name can be hosted in Europe; a generic top-level domain can serve a local organisation. Counting names by suffix is not the same as measuring national dependency.

Replication and complementary vantage points can test geographic sensitivity. Where results differ, the difference is data rather than an inconvenience to average away. It may reveal anycast policy, content localisation or blocking.

The project’s public-interest role is strongest when coverage gaps are mapped explicitly. Regions with weaker source agreements should not disappear into a global percentage. A historical record can reduce knowledge inequality only if users know where its instrument was less able to see.

Billions of queries matter only when their context survives

The measurement pipeline turns target lists into scheduled queries. Workers issue requests for defined record types, receive responses, normalise fields and store observations with timestamps and metadata. At the scale reported by OpenINTEL—about 5.9 billion data points per day—the operational challenge is not sending one DNS packet. It is completing the daily cycle reliably without overwhelming authoritative infrastructure or losing the conditions behind the result.

Workers need rate control, retry policy and clear source identification. A timeout can mean no service, packet loss, rate limiting, a transient failure or an intentional block. Retrying too aggressively can create the harm the project is trying to avoid. Giving operators a public explanation and contact path makes the traffic accountable.

Responses must be parsed across many record types and DNS edge cases. Names can contain unusual encodings. Delegations can be lame or cyclic. DNSSEC adds signatures, keys and denial-of-existence records. Truncation can shift a query from UDP to TCP. Authoritative servers can return different answers by source location. Normalisation must preserve meaning without turning every packet into an unmanageable format.

The system also needs a definition of completion. A daily run may finish for most targets and miss a subset. Storing only successful responses would make absence invisible. Researchers need to know which queries were attempted, which failed and whether a platform outage affected a period. A missing record should not automatically become evidence that a domain removed it.

OpenINTEL’s homepage reports 13.6 trillion cumulative data points since 2015. That figure communicates scale and remains project-reported. Its analytical value depends on how “data point” is defined across products and time. A count can grow through more domains, more record types or more frequent observations. Users should consult the relevant dataset methodology rather than compare cumulative totals as a simple measure of DNS growth.

At this volume, engineering decisions shape research. Partitioning, compression, indexes and storage formats determine which queries are practical. Data migrations can change representation. Quality-control systems need to identify incomplete runs. The archive is a scientific instrument and a data platform at once.

Forward DNS reveals configuration rather than application behaviour

Forward measurement begins with a name and asks for selected records. Delegation data can show which authoritative providers serve the domain. Address records can show hosting or CDN relationships. Mail exchange records expose email infrastructure. DNSSEC records indicate deployment and algorithm choices. Other types reveal service and policy configuration.

Repeated observations make transitions visible. A domain can move from one authoritative provider to another, add IPv6, enable DNSSEC or change mail service. At large scale, researchers can estimate adoption and concentration. They can examine whether changes occur gradually or around an event.

A returned record remains an observation at a time and vantage. An A or AAAA address does not prove that a website responded, that the address served the same content to users or that the application was secure. A mail exchange does not prove successful delivery. A DNSSEC signature can be present while validation fails elsewhere. Application-layer testing requires separate methods.

CDNs and anycast complicate interpretation. An authoritative or recursive response can vary by source location. A domain may return addresses selected for the OpenINTEL vantage rather than the addresses a user in another region receives. Split-horizon systems deliberately give different answers to internal and external clients. The project’s measurement is not false; it is one view.

Caching adds another distinction. OpenINTEL’s active system may query through defined resolver paths or authoritative infrastructure depending on the dataset. It does not observe what every recursive resolver has cached. A user can receive an earlier value until TTL expiry. Changes in the archive can precede or follow user-visible transitions.

The platform’s value is strongest where the research question matches the method: how configured DNS answers observed by the project changed. It weakens when the answer is used as a proxy for popularity, application success or user experience without additional evidence.

Reverse DNS and address-space products connect naming with network administration

Reverse DNS starts from an IP address and asks which name, if any, is associated through the in-addr.arpa or ip6.arpa hierarchy. OpenINTEL expanded into IPv4 reverse measurement, creating another large longitudinal view. The dataset can reveal administrative naming patterns, infrastructure changes and the presence of records across address space.

A PTR record is not authoritative proof of who uses an address or what service it provides. Address holders can leave records stale, use generic names or delegate reverse zones. Cloud and access networks may apply systematic naming. Some addresses have no reverse entry. The data is useful for classification and change, not a universal identity map.

Measuring IPv4 reverse space is tractable relative to IPv6 because the address population is smaller and enumerable under defined policy. IPv6 is too large for exhaustive address-by-address scanning. Research must use allocated prefixes, observed addresses or other target-selection methods. This difference prevents the IPv4 methodology from being projected onto the newer protocol without qualification.

The project also publishes RIR-level and network-prefix products. Temporal prefix top lists attempt to rank network prefixes under defined observations. Such lists can support measurement sampling and research, but they are not an objective hierarchy of network importance. A prefix can appear prominent because of the list construction and service population. Business value, traffic and user count remain separate.

These products broaden OpenINTEL from a domain observatory toward a naming-and-addressing platform. They also increase the need for careful labels. “DNS history” can cover domain records, reverse names, certificate-derived targets and inferred zone changes, each with a different population and cadence.

The expansion is analytically valuable because internet infrastructure joins names, addresses and networks. A hosting migration may appear in forward records and prefix relationships. Reverse naming can provide operational context. RIR data can group observations. The connection remains an inference framework rather than a complete ownership ledger.

Zonestream narrows the gap between daily snapshots and intra-day change

A daily scan records one broad state. The DNS can change several times between scans. A malicious campaign can activate and disappear. A large provider can migrate records in stages. A configuration error can be introduced and corrected before the next scheduled run.

Zonestream and related work aim to provide more event-oriented evidence by inferring zone changes and producing feeds closer to the time of change. The approach complements the daily archive rather than replacing it. A stream can identify rapid transitions; the daily pipeline supplies broad, consistent snapshots.

Faster measurement creates load and interpretation challenges. More frequent queries increase traffic to authoritative operators. Changes in a zone source do not always mean the same thing as changes in observed DNS answers. A stream can contain bursts from maintenance or automated systems. Consumers need to distinguish raw events from meaningful infrastructure transitions.

The product shows how OpenINTEL’s architecture has evolved beyond one batch process. The project began by proving that a huge daily run could complete. It later added ways to observe change at different temporal resolution. This broadens use cases while complicating comparability.

Researchers should state which cadence supports a claim. A daily dataset can show that a configuration differed between dates. A stream can show an intermediate sequence. Neither necessarily explains why the operator acted. Combining them with registry, certificate or incident evidence can strengthen the account.

Zonestream also increases the operational importance of continuous availability. A missing daily scan creates one gap. A failed event feed can lose a sequence that is difficult to reconstruct. Redundancy, replay and monitoring become part of the research method.

Storage is the project’s most durable output and largest obligation

OpenINTEL’s archive is valuable because yesterday’s DNS cannot be queried directly. Once a state changes and caches expire, repeated measurement may be the only public evidence that the project’s vantage observed it. That makes storage and data integrity central infrastructure.

A decade-scale archive needs more than capacity. It needs versioned schemas, checksums, replication, documented migrations and a way to preserve the relationship between observations and methods. A column renamed without a migration record can break reproducibility. A compression change can improve cost and complicate older tools. A corrupted partition can remove evidence that no new scan can regenerate.

Retention creates financial pressure. Billions of daily points require compute, network and storage. Public sources do not disclose a consolidated annual cost. The burden is shared among partner institutions and funding programmes. As the archive grows, the project must choose between retaining raw detail, producing derived datasets and controlling access.

Query cost is another constraint. A researcher may want to scan years of records across millions of domains. Allowing unrestricted queries can overwhelm the platform. Download products and controlled access distribute work but require users to store and process data themselves. Cloud-hosted copies could improve access while creating cost and governance questions.

Longitudinal integrity also depends on preserving absence. A day with no record can mean the domain lacked the value, the query failed, the target was not in the list or the pipeline was incomplete. The archive needs enough control data to distinguish those states. Otherwise, a trend can be an artefact of instrumentation.

The project’s future will be determined partly by whether institutions continue to fund this invisible work. New record types and dashboards attract attention. Maintaining old bytes, documentation and context creates the historical asset. If storage or specialist staff are lost, the archive’s continuity cannot be bought back later.

Four institutions share the platform without one visible budget

OpenINTEL’s organisational resilience comes from partnership. The University of Twente supplies academic leadership and researchers. SIDN contributes registry knowledge and support. NLnet Labs brings DNS-software and operational expertise. SURF contributes national research infrastructure. The combination is stronger than a project dependent on one principal investigator and one grant.

The model is also opaque in conventional financial terms. There is no consolidated project revenue, expense or staff count. Hardware may be funded by one partner, researchers employed by another and network capacity supplied by a third. Public pages identify roles but do not provide one voting charter or asset register.

That does not make the project ungoverned. Decisions occur through institutional relationships and a core team. It does mean that changes in partner priorities can affect the platform without appearing as a corporate event. A grant ends, a server reaches replacement age or a specialist moves roles. The archive can continue while development capacity narrows.

The concentration of expertise is a risk. Long-running measurement systems accumulate knowledge about quirks, data migrations and operator relationships. Documentation and succession are as important as new code. A partnership can spread the burden only if more than one institution can operate critical components.

The model also shapes accountability to measured operators. A visible project identity and abuse contact allow authoritative providers to report excessive traffic. Partners with DNS community standing can negotiate problems. The system needs to preserve that trust as scale and products expand.

OpenINTEL is best described as shared research infrastructure. Its authority comes from the quality and continuity of its evidence, not a statutory claim over the DNS. It measures a distributed system under the tolerance of organisations that remain free to block or restrict it.

Open data still depends on contracts, licences and long-term funding

OpenINTEL promotes research access and releases eligible datasets under CC BY-NC-SA 4.0. The licence requires attribution, restricts commercial use and applies share-alike conditions. Other material remains controlled because zone-access agreements or source contracts do not permit unrestricted redistribution.

This arrangement can disappoint users who hear “OpenINTEL” and assume that every observation is freely downloadable for any purpose. The project name describes a research commitment, not ownership of every input. A registry can allow measurement under conditions without granting the right to republish its complete zone-derived data.

The non-commercial restriction supports academic sharing and limits some industry reuse. A company may need a separate agreement for a commercial product. The project does not publish a universal commercial licence or price. Users should contact the operators rather than assume that access terms can be inferred from the open datasets.

Controlled data can still support research through applications, institutional agreements or derived outputs. The process introduces selection and administrative cost. Researchers with established affiliations may gain access more easily than independent analysts. Reproducibility is harder when the underlying dataset cannot be redistributed.

The tension is structural. Longitudinal DNS research benefits from broad source access. Registries and operators have contractual, security and commercial concerns. A platform that violated those agreements might publish more in the short term and lose future access. Sustainable openness sometimes requires a documented boundary rather than maximal release.

For data users, the correct practice is dataset-specific. State the source, licence, coverage and access condition. Do not describe controlled data as public. Do not assume an open derived table contains the full underlying archive. The governance of the data is part of the method.

The current open datasets use a non-commercial Creative Commons licence, while other material remains controlled by source contracts. This supports academic work and prevents a simple assumption that all archive content is free for any business use.

Commercial analysts may want historical DNS evidence for security, market research or due diligence. Their demand could help fund infrastructure and can conflict with restrictions imposed by registries and the expectations of measured operators. A paid path would need to distinguish service and support from rights the project does not possess.

The partners could provide derived aggregates, controlled research environments or negotiated licences for eligible data. Each model changes who can reproduce a result. A private product built from a public-interest archive may generate value without returning methods or fixes.

The governance question is not whether commercial use is good or bad. It is whether revenue arrangements preserve the longitudinal record, honour source rights and avoid making the most complete data available only to well-funded users.

OpenINTEL’s sustainability may eventually require more formal access tiers. The credibility of those tiers will depend on transparent criteria and a protected public baseline. The archive became valuable through shared research. Financing its future should not make its past impossible to examine.

A DNS answer records configuration, not intent or harm

Large DNS datasets invite categorical conclusions. A record points to an address associated with a provider, so the domain is said to be hosted there. A name returns NXDOMAIN, so it is declared nonexistent. A certificate-derived name resolves, so the service is assumed active. Each inference can be useful and wrong in a particular case.

A DNS response records what the queried infrastructure returned under the measurement conditions. It does not show why the operator configured it. The address may be a redirect, sinkhole, parked page or shared CDN edge. The application may reject the hostname. The record may be stale. A transient server failure or rate limit can produce an apparent absence.

NXDOMAIN means that the responding DNS path asserted nonexistence for the queried name under its current state. It does not prove the name never existed or will not exist later. Delegation errors and inconsistent authoritative servers can create varying results. Repeated observations and direct authoritative checks improve confidence.

Security analysis requires even more caution. Rapid domain or address changes can be associated with abuse and with legitimate CDNs, failover and migrations. A record does not prove that a domain is malicious. Labelling requires additional evidence such as content, campaign relationships, registration and observed behaviour.

User demand is entirely outside active measurement. OpenINTEL generates its own queries. It does not see how often users request a name, what recursive caches serve or which answer produces traffic. Passive DNS or resolver telemetry addresses different questions and carries different privacy concerns.

The project’s strongest analytical culture is one that treats these boundaries as first-class. A huge archive can support better inference because patterns and history are available. It does not change the logical status of one observation. The measurement remains an input to explanation, not the explanation itself.

An NXDOMAIN response can indicate that a name does not exist in the relevant DNS view. A timeout can indicate an unreachable server, rate limiting, packet loss or deliberate refusal. A SERVFAIL can arise from validation, delegation or transient operational problems.

Longitudinal analysis should preserve these categories instead of collapsing them into “down.” A domain that moves from a valid answer to NXDOMAIN has a different history from one that times out intermittently. A parser or retry change can alter the measured distribution without any operator changing configuration.

The distinction is particularly important in security and policy research. A non-response is not proof that a domain was removed or censored. Additional vantage points, authoritative queries and application checks may be needed.

OpenINTEL’s scale makes error classification consequential. A small methodological choice can affect millions of records. Careful treatment of negative evidence is one of the clearest ways the project can prevent a vast archive from producing overconfident conclusions.

Longitudinal claims depend on method history and missing observations

A decade-scale dataset contains two histories: the history of the DNS and the history of the instrument. Hardware is replaced, query software changes, parsers are fixed and source agreements expand. A record type added in 2024 cannot be compared directly with an absence in 2018. A new retry rule can improve completion while changing the probability that a slow server appears responsive.

For this reason, method versioning is part of the data. Each observation series should be tied to the target list, query type, measurement vantage, software release and known operational events. Missing days need explicit flags. An analyst should not infer that millions of domains changed when a worker cluster failed or an input feed arrived late.

The problem becomes harder after storage migration. A new schema may compress repeated fields, normalise names differently or reclassify errors. Those changes can make the archive cheaper and easier to use while altering old queries. Preserving raw or sufficiently detailed source records, migration code and validation samples allows researchers to distinguish a DNS trend from a database transformation.

Reproducibility does not require the entire platform to remain frozen. It requires enough evidence to reconstruct how a result was produced. Published work should identify dataset release or access date, population filters and code. When contract terms prevent redistribution of raw records, researchers can still publish methods, aggregates and validation checks within the agreement.

OpenINTEL’s partnership with external researchers and replication work is important here. A second implementation or independent measurement will not produce identical answers, but differences can expose assumptions. Replication is particularly valuable where anycast, blocking or list licensing makes one vantage structurally incomplete.

The archive’s age increases the cost of a silent change. A minor parsing fix applied retrospectively can alter years of data. Leaving the error untouched can perpetuate a known flaw. The responsible approach is to document the correction, preserve the original state where possible and specify which version underlies a result.

This is the unglamorous work that separates infrastructure from a collection of files. The platform’s reported trillions of data points matter only if future users can tell which points belong in the same comparison. Longitudinal science is an exercise in preserving context at the same scale as observations.

Responsible scanning must remain visible to the operators bearing its cost

Active measurement consumes resources outside the project. A DNS query is small, but billions of queries reach authoritative systems that vary from large anycast platforms to modest servers. Rate controls and scheduling reduce impact; they do not make the cost zero. The ethical basis of the work therefore includes transparency and a practical path for operators to object.

A responsible platform uses identifiable source addresses, publishes a description of its traffic and monitors contact channels. It should honour justified requests to reduce or block measurement and investigate reports of unusual load. Those measures do not create universal consent. They make the project accountable for an activity that is technically possible without prior permission.

The burden is not evenly distributed. A heavily delegated provider may receive queries for millions of names, while a small operator sees only a few. Some servers are configured to rate-limit unfamiliar traffic, causing an apparent measurement failure. Others may return deliberately generic answers. An analyst needs to recognise that operator defence changes the observed dataset.

Privacy risk is different from passive resolver logging. OpenINTEL generates queries from target lists and does not watch individual users. That greatly limits exposure to user behaviour. The archive can still contain names that identify organisations, devices or services, including subdomains derived from certificate logs. Publishing detailed historical records can make forgotten infrastructure easier to find.

Access controls and licensing can reduce misuse without turning every DNS observation into confidential data. The appropriate boundary depends on the source, granularity and risk. Broad aggregates and research datasets may be safe to release, while raw zone-derived records remain contractually controlled. A project committed to open science has to explain these distinctions rather than present access as all or nothing.

Ethical review must also account for downstream claims. A research paper that labels a domain or country insecure can cause reputational harm when the measurement actually captured a transient error. The project cannot police every user, but clear caveats, terms and examples can shape better practice.

OpenINTEL’s legitimacy rests partly on the fact that measured operators can see who is asking. That visibility should remain a design requirement as products become faster and more varied. An observatory earns the tolerance of the system it observes by making its own behaviour open to inspection.

Active DNS, passive DNS and general scanning answer different questions

OpenINTEL is sometimes compared with passive DNS databases, internet-wide scanners and distributed probe systems. The comparison is useful only after the observation model is separated.

Passive DNS collects records seen in real query traffic at recursive resolvers or other observation points. It can reveal what users or systems requested and which answers were returned through those paths. Coverage depends on participating sensors and raises privacy and contractual issues. A passive database may see a popular malicious domain quickly while missing a quiet domain that no observed user asks for.

OpenINTEL chooses its targets and asks the questions itself. It can measure the same population every day even when no user visits the names. That regularity supports longitudinal comparison. It cannot infer popularity or cache behaviour from the active queries. The two methods are complementary: one reflects observed demand at selected resolvers, the other reflects configured answers for selected targets from measurement vantage points.

A scanner such as ZMap begins with addresses and asks whether a service responds, often followed by a protocol handshake. It can map exposed services without relying on domain lists. DNS measurement can identify names and delegations that point to shared infrastructure, including records whose services are unreachable. Again, the methods see different surfaces.

A distributed probe platform such as RIPE Atlas can ask DNS questions from many networks and locations, revealing geographic and resolver-dependent differences. OpenINTEL emphasises breadth, repetition and archive scale from its measurement infrastructure. A smaller population from many vantage points can answer a question that one vast daily scan cannot.

These distinctions prevent a common mistake: treating all large internet datasets as interchangeable evidence. A domain absent from active DNS, unseen in passive data and unreachable in an address scan may still exist behind access control or split-horizon naming. A domain observed by all three methods has stronger evidence of public operation, but even that does not establish who used it or why.

The strategic opportunity is not to build one universal database. It is to connect methods through explicit timestamps, populations and uncertainty. OpenINTEL can provide the longitudinal naming layer in that larger evidence system. Its value grows when analysts know which questions require another instrument.

DNSSEC adoption becomes legible only through repeated measurement

DNSSEC is a useful example of why a daily archive matters. A zone can publish delegation material, authoritative servers can serve signed records and validators can decide whether the resulting chain is secure. Those stages do not necessarily move together. A country-code registry may enable signed delegations while many domain holders remain unsigned. A domain may publish keys but fail validation because signatures expire or a parent record is wrong. A one-day scan can count states; it cannot show how operators entered, left or repaired them.

OpenINTEL can follow the appearance and disappearance of relevant records across a defined population. Researchers can separate domains that remain unsigned from those that are configured intermittently, identify changes around algorithm or key transitions and measure how long broken states persist. The repeated method turns an adoption percentage into a set of operational paths.

The interpretation still depends on the measurement design. Seeing DNSKEY or DS records is not identical to performing every validation step exactly as a user’s recursive resolver would. Responses can differ by vantage, and a domain can be signed while its application remains unavailable. A project studying DNSSEC needs to state which records, validation logic and error categories it used. Comparisons across years need to account for list growth and software changes.

Those qualifications make the result more useful, not less. Standards discussions often rely on headline adoption figures that hide the cost of maintenance. Longitudinal data can show whether a change produces durable configuration, a burst of experimentation or a recurring failure pattern. It can distinguish a slow deployment from one that reached a plateau because the remaining population has different incentives.

The same logic applies to IPv6 records, mail security and other infrastructure practices. A technology is not deployed merely because one record appeared once. It becomes infrastructure when configuration persists, failures are repaired and dependent systems behave consistently. OpenINTEL’s archive can reveal those transitions because it preserves enough daily evidence to separate an event from a habit.

For operators, that history can be uncomfortable. A long record makes repeated misconfiguration visible. It can also establish that a problem was corrected before an incident or policy intervention. Measurement does not assign intent, but it changes the quality of the argument. The discussion moves from recollection to dated evidence.

Provider concentration appears in patterns, while causation stays outside the dataset

The DNS can expose where infrastructure is concentrated. Many domains may delegate to the same authoritative provider, point mail to the same service or resolve web names into address space associated with a small set of platforms. Tracking those relationships over time can show consolidation, migration and dependence that would be difficult to reconstruct after a market shift.

OpenINTEL is well suited to the descriptive part of this work. Repeated records can reveal that a provider’s nameserver domains appear across a growing share of a measured population, that an acquisition is followed by a change in naming patterns or that domains move away after an outage. Address and autonomous-system enrichment can connect names to network infrastructure, subject to the accuracy and timing of routing data.

The archive cannot explain every reason for a move. A domain may use a provider because of price, performance, security, a registrar bundle or an organisational merger. Shared nameserver branding may conceal several independent infrastructures. A large platform can use many autonomous systems, while one autonomous system can host unrelated customers. DNS data supports a concentration measure; it does not prove market power or customer satisfaction.

The policy consequence is direct. A regulator may be tempted to treat a high measured share as evidence of harmful control. An operator may treat the same share as evidence that a service has earned trust. The dataset can establish the pattern and the date. Economic and legal conclusions require contracts, ownership records, outage evidence and user switching costs.

Longitudinal measurement adds two insights that a market snapshot misses. First, it can show the speed of concentration. A slow decade-long move has different causes and remedies from a sudden change caused by a product withdrawal. Second, it can show reversibility. If domains frequently switch providers, a concentrated market may still have practical mobility. If delegations persist despite repeated failures, lock-in may be deeper than the headline share suggests.

The history can also reveal hidden common-mode risk. Organisations may believe they use diverse application hosts while delegating DNS, mail and certificates through the same provider family. A failure in that shared layer can affect otherwise separate systems. OpenINTEL does not model the entire dependency chain, but it can provide the naming evidence from which a more complete map begins.

The most responsible use of the archive is therefore to treat concentration as an observed structure and then investigate the mechanism. A measurement platform should make dependence visible without pretending that a list of records contains the whole political economy of the internet.

Certificates, mail records and names can be joined without becoming one system

DNS does not operate in isolation. Certificate Transparency logs expose names submitted for public certificates. Mail records identify exchange infrastructure and policy. Address records point toward hosting networks. Reverse names can provide administrative clues. OpenINTEL’s source products and repeated queries make it possible to examine how those systems line up over time.

Certificate-derived names expand the observable population beyond apex domains available in a zone list. They can reveal service subdomains and temporary names that matter to PKI research. The source is selective: it favours publicly logged certificates and includes names that may never have served traffic. A certificate can remain in a log after the associated service disappears. Active DNS measurement adds a current configuration observation, not proof that the certificate is installed or trusted by a browser.

Mail records create a different map. MX targets and related TXT records can show movement toward hosted mail providers, adoption of authentication policies and configuration errors. Repeated observations can identify whether a policy was maintained or briefly tested. They do not show message volume, delivery success or how receiving systems applied the policy.

Infrastructure records can also support incident reconstruction. If a domain changed nameservers, mail hosts and addresses within a narrow period, investigators gain a timeline. The sequence may reflect a legitimate migration, recovery from compromise or takeover. DNS history narrows the questions; it does not label the event.

The value comes from joining evidence while preserving its origin. A certificate log, a zone, an active response and a routing table each have different timing and authority. Combining them in one analysis can reveal a relationship, but the conclusion is only as strong as the weakest join. A name reused by several services, an address behind a shared platform or a stale certificate can create a false association.

OpenINTEL’s role is strongest when it keeps these datasets distinguishable. Researchers should be able to say that a name entered the target list through a certificate log, returned a particular record during the active measurement and mapped to a network prefix under a separately dated routing view. That sentence is less dramatic than saying the platform “knows the internet,” but it is reproducible.

This multi-source discipline is also a defence against retrospective certainty. After an incident, analysts naturally search for a story in the data. A versioned measurement trail forces the story to respect what was observable at the time. It can show that two events were correlated without claiming that one caused the other.

Prefix rankings extend the archive from names toward networks

The project’s temporal prefix top lists illustrate how OpenINTEL can produce derived products rather than only raw DNS answers. A prefix ranking aggregates evidence at a network level and tracks how prominence changes over time. This can help researchers select representative networks, study infrastructure concentration or compare measurement targets without rebuilding the same joins for every paper.

A prefix is not an organisation. Routing announcements can change, address space can be leased, and one network can host many unrelated services. Rankings depend on the names and record types included, the date of the routing data and the metric used to count prominence. A high position can mean broad hosting, shared infrastructure or a list-selection effect.

The temporal design is the useful part. A static list quickly becomes stale as services move and routing changes. Repeated rankings can show when a network enters or leaves a prominent position and allow researchers to reproduce a population appropriate to an earlier date. They also expose instability that a single “top networks” list would conceal.

Derived datasets lower the cost of research. They can make sophisticated infrastructure analysis available to teams without the storage and compute needed to process the full archive. That benefit increases the project’s responsibility to publish methodology and versioning. Users may treat a convenient ranking as objective ground truth when it is one view produced from selected inputs.

The expansion toward network-level products does not change OpenINTEL into a routing observatory or RIR. It uses public registry and routing context to enrich DNS observations. Route validity, ownership disputes and operational performance remain separate questions.

This boundary is strategically healthy. The project can offer common derived evidence without claiming authority over the entities it ranks. The product is most valuable when it saves computational work while leaving the analytical judgement visible.

Security research gains a timeline rather than a verdict on maliciousness

Historical DNS data is attractive to security teams because abuse infrastructure often moves. A domain may rotate addresses, nameservers or mail systems; a campaign may reuse providers; an incident may be discovered after the relevant configuration has changed. OpenINTEL can supply dated observations that make those transitions recoverable.

That evidence is especially useful for scoping. Investigators can ask when a suspicious domain first appeared in a measured list, whether its authoritative service changed around an event and which other names shared the same infrastructure at the time. A security researcher can use those relationships to generate hypotheses and choose which systems require closer examination.

None of the records establishes maliciousness by itself. Fast changes can be a sign of evasion, but they also occur in content-delivery networks, disaster recovery and legitimate migrations. Shared hosting puts benign and harmful domains on the same address. A nameserver association can reflect a registrar default rather than common control. Even a pattern that strongly resembles a known campaign needs corroboration from content, registration, malware, telemetry or legal evidence.

That distinction matters because historical datasets can make associations look more durable than they were. A domain that shared an address for one day may be grouped with another for years in an analyst’s derived database. Time windows and confidence should travel with the relationship. So should the source population: a certificate-derived subdomain and a registry-zone apex name did not enter the archive in the same way.

OpenINTEL’s role is to preserve the infrastructure facts that may otherwise disappear. Labelling, attribution and response belong to separate processes. Security work is strongest when the archive narrows uncertainty without being asked to resolve intent.

The archive makes infrastructure change contestable

Every day of measurement adds operational value and storage obligation. The project must maintain query capacity, replace hardware, migrate databases and keep specialists who understand both DNS and the history of the system. None of those tasks is secured merely because the archive has become important.

The four-partner model spreads risk, but it also makes sustainability harder to read from outside. One institution may fund staff, another compute and another access to source data. A change in any component can reduce coverage without a conventional company announcement. Public financial accounts for the project as a whole are not available, so users should watch technical and institutional signals rather than assume continuity.

Archive stewardship requires redundancy beyond backups. A replicated copy must include schemas, method records, source-list versions and the knowledge needed to interpret anomalies. A pile of files at another site is not a functioning historical instrument. Collaboration with organisations such as CAIDA can improve resilience and methodological comparison, although contractual data may limit what can be copied.

The licensing model also shapes the future. Non-commercial open access supports academic reuse but restricts some commercial applications. Controlled zone data cannot simply be released under a broader licence. The partners may need funding arrangements that preserve research access while recovering the real cost of storage and support. A commercial pathway, if developed, should not silently narrow the public record on which the project’s legitimacy rests.

Success can itself create fragility. As more papers and policy claims depend on OpenINTEL, a missing period or changed dataset affects a wider community. The project may need more formal release records, service expectations and preservation policy than a research system normally publishes. These are not signs that it should become a company. They are signs that it has become infrastructure.

The most valuable future milestone may be one that produces no headline: a transparent migration to new storage with the old results, caveats and access paths intact. OpenINTEL has already shown that it can measure at remarkable scale. The harder test is whether the organisations behind it can preserve the conditions that make ten years of measurements comparable to the next ten.

OpenINTEL cannot tell the whole history of the DNS. It can show that a defined name returned a defined answer from its vantage on a date, and it can repeat that observation across enormous populations. That is enough to transform many claims.

A provider can say that DNSSEC adoption increased; the archive can show the measured transition and population. A researcher can claim that authoritative hosting concentrated; the data can reveal which lists and periods support it. An incident investigator can reconstruct when a record changed. Another analyst can challenge the methodology rather than accept a screenshot.

The project’s scale is striking: hundreds of millions of domains, billions of daily points and trillions of cumulative observations according to its own current metrics. The more important achievement is continuity across institutional and technical change. The archive makes the DNS’s past available as evidence rather than memory.

That evidence remains bounded by lists, contracts, vantage and method. OpenINTEL’s credibility comes from preserving those boundaries. Calling it a complete copy of the DNS would overstate the project and weaken the usefulness of its real record.

A historical observatory does not need to see everything. It needs to state what it saw, preserve the conditions and remain available long enough for change to be measured. OpenINTEL has built that kind of instrument for naming infrastructure. Its next challenge is ensuring that the archive and the institutions behind it remain as durable as the trends researchers hope to study.