Summary

  • ICANN’s final data overview says an additional AI-assisted analysis after Public Comment identified 1,612 Han-script similarity cases. After further consultation with script experts, the cases were incorporated into the Chinese, Japanese and Korean String Similarity Evaluation data.
  • Those additions enlarge a frozen input, not a body of final findings. The SSE tool proposes potential contention sets; the independent panel may add, remove or adjust them, must give reasons, and faces a 21-day applicant challenge route for factual, procedural or system error.

The consequential number in ICANN’s 30 July release is not the number of applications. It is 1,612: the count of additional Han-script similarity cases that entered the final data after the public consultation had closed.

The final data overview explains how they were found. The Chinese script expert conducted further analysis using AI-assisted factoring based on four methods: convolutional neural networks, stroke-structure analysis, Zernike moments and grid histograms. ICANN says the work produced 1,612 additional Han cases. After further consultation with script experts, the cases were included under a conservative principle in the Chinese, Japanese and Korean data files.

That disclosure matters. A code-point or sequence mapping can widen the group of string pairs sent forward for examination. A new mapping may therefore create work for a panel and uncertainty for an applicant. But the disclosure does not establish 1,612 application conflicts, rejected strings, errors or final contention relationships. It records an expansion of the evidentiary field.

The governing distinction is simple: better search is not delegated judgment.

What changed between the commented and frozen data

The October 2025 material put one Common Similarity Data file and 26 script-specific files out for Public Comment. It already described a conservative design. Where a pair was ambiguous, the data could include it so that the panel, rather than an early filter, would have the chance to review it. Transitive relationships could also place pairs in a common set even when experts had not directly classified that exact pair as confusing.

The consultation ran from 16 October to 4 December and drew 11 submissions. The summary records requests to add or remove particular pairs, provide public access to the SSE tool, maintain the data over time and use conservative categories where doubt remained. ICANN answered specific comments, said it would discuss the feasibility of public tool access, and stated that the finalized data would be used for the 2026 round while later input would be queued for future work.

The July 2026 final overview then disclosed the additional Han analysis. The public record is therefore not silent about the increment. It names the methods, the count, the affected script files and further expert consultation.

What the checked public materials do not provide is a clean, machine-readable draft-to-final change file isolating every one of the 1,612 additions, its prior state, its provenance class and its expert disposition. Nor do they show which added mapping will actually affect an applied string. That is a visibility limit, not evidence that no internal review or change history exists.

The distinction matters because “after Public Comment” is not the same as “outside scrutiny.” Finalization normally incorporates work after a consultation. A new consultation is not automatically required every time an expert corrects or expands a dataset. The proper governance question is narrower: can a later reader tell what materially changed, why it changed, who reviewed it and how it affected a consequential result?

The frozen version creates equality and responsibility

The final data and guidelines are both version 1.0 and dated 23 July 2026. The guidelines identify those versions as the ones applicable to the 2026 round. New input may still be sent to ICANN, recorded and discussed with relevant experts, but it will not change the data or guidelines used in this round.

That freeze solves a real problem. Applicants should not be compared under a moving set of similarity rules. A pair considered distinct in September should not become confusing in November merely because a file changed without a visible round boundary. Reproducibility requires one known input version.

Freezing the file does not freeze judgment. The SSE tool uses the data to generate a pre-screening report of potential contention sets. The panel takes that report as one input and may consider additional analysis. It can adjust, add or remove potential sets. It makes the final determination and must provide relevant reasons for each decision.

The guidelines make the handoff explicit. Code points and sequences receive similarity levels in the data. Whole strings receive the panel’s judgment. When the panel departs from the tool output, its rationale must explain the deviation. Where a case is difficult, the guidelines favour the conservative course of treating the strings as similar. That bias belongs to the disclosed evaluation design; it is not proof of the outcome in any unexamined case.

The resulting authority chain has four distinct states:

  1. a similarity case exists in a script data file;
  2. the tool uses the frozen data to suggest a potential contention set;
  3. the panel decides the whole-string outcome and gives a rationale; and
  4. a timely challenge may test whether factual, procedural or system error affected that outcome.

Collapsing those states would produce two opposite mistakes. One would accuse an algorithm of rejecting an application merely because a mapping entered the data. The other would treat a human panel label as unreviewable because a tool found the pair first. Neither follows the published process.

The consequence belongs to the panel decision

The Applicant Guidebook gives the final evaluation practical force. Depending on what the string is compared with, a finding of visual similarity can stop an application, put it on hold or form a contention set with another application. An outcome applies across the relevant variant-string set. ICANN says outcomes and rationales will be published on the program’s Evaluation Results Page.

That consequence is why attribution matters. The data’s job is to enlarge or narrow the search space. The tool’s job is to make the comparison volume manageable. The panel’s job is to decide whether complete strings are so visually similar that coexistence would create a probability of user confusion. ICANN’s job is to apply the program consequence. An applicant’s challenge right tests a defined class of error; it does not turn every disagreement into a new merits proceeding.

The current rules give an applicant 21 days after receiving the result to file an Evaluation Challenge. The grounds are bounded: factual, procedural or system error. If an error is confirmed, the evaluation is performed again with the finding in mind. If no error is found, the original outcome stands. New material that would substantially change the application is not a substitute for identifying an error in the evaluation.

That architecture is more accountable than either pure automation or unstructured discretion. Its weakness would arise if the layers could no longer be joined. A published outcome without the frozen input identity would be hard to reproduce. A tool result without panel reasons would be a suggestion dressed as authority. A rationale without a correction history would obscure whether a challenge changed the result.

A change-to-decision receipt

The public answer is not to publish confidential applications before their authorized release, expose protected panel deliberations or turn every expert disagreement into a spectacle. It is to keep a compact chain from material data change to institutional result.

Daniel Kade’s proposed change-to-decision receipt begins with the data version and fingerprints of the affected files. For each material draft-to-final addition, removal or reclassification, it records the relevant code point or sequence, the before-and-after state and a provenance class: Public Comment response, expert correction, expert expansion, AI-assisted candidate, transitivity injection or mechanical transformation.

The receipt then records the input date, method and supporting reason; the affected Chinese, Japanese, Korean or other files; the script-expert consultation state; and whether the mapping is used for screening or retained only for documentation. It does not claim unanimous endorsement unless the record proves it.

If the mapping later contributes to a potential contention set, the protected record links the tool version and that contribution. The public record, at the stage permitted by the program, then links the panel’s whole-string decision, the guideline applied and the reason for adopting, adding, removing or overriding the tool suggestion. Any challenge, error finding, reevaluation, correction and supersession stays attached.

This receipt is not an announced ICANN requirement. It is a proportionate way to preserve three things the final release already treats as separate: data provenance, computational assistance and accountable judgment.

ICANN already publishes the final data in human-readable and machine-readable form. A draft-to-final delta and aggregate counts by script and provenance class would make that transparency more useful. They would let applicants and script communities inspect material movement without converting preliminary application information into a public case file.

The 1,612 additions should therefore be read neither as a scandal nor as a certificate of correctness. They show why the process needs both scale and restraint. Automation can discover more candidates than unaided manual comparison. Script experts can supply knowledge no generic model possesses. A panel can judge the whole label in context. A challenge can correct a defined error. Each layer is useful because it does not pretend to be the next one.

Sources

  1. ICANN — String Similarity Evaluation data and guidelines announcement, 30 July 2026
  2. ICANN — final String Similarity Evaluation Data overview, version 1.0
  3. ICANN — final String Similarity Evaluation Guidelines, version 1.0
  4. ICANN — String Similarity Evaluation resource page
  5. ICANN — Public Comment proceeding on the SSE data
  6. ICANN — Public Comment Summary Report on the SSE data
  7. ICANN — draft SSE Data overview published for Public Comment
  8. ICANN — New gTLD Program: 2026 Round Applicant Guidebook
  9. ICANN — FAQ on challenges to String Similarity Evaluation outcomes