Summary

  • Matcha moves privacy-label work closer to the code and SDKs that produce data practices, while keeping developers responsible for interpreting its suggestions.
  • A study found accuracy gains for most of 12 participants, but the tool took longer than the store form and the results do not certify labels at platform scale.

App-store privacy disclosures appear beside the download decision, yet developers usually complete them in a separate form. That separation matters. A form asks people to classify what an app collects, shares and uses; the evidence may be spread across first-party code, third-party libraries, configuration and server-side handling. The person filling in the form can know the app’s intention and still miss what a dependency does.

Apple and Google have built different systems around this problem. Apple’s App Privacy responses are supplied at app level, must include the practices of third-party partners and should be updated when practices change. Google Play’s Data safety section is a global declaration for a package, covering currently distributed versions and relevant SDK behavior. Google says developers alone have the information needed to make complete and accurate declarations; its review is not designed to verify their accuracy. These are separate taxonomies and workflows, not interchangeable labels. (Apple; Google Play)

Put the evidence where the code lives

Lorrie Cranor has worked across privacy engineering, public policy and usable security. Carnegie Mellon’s CUPS lab says its privacy nutrition-label design team was led by Patrick Gage Kelley and included Cranor; the project aimed to make privacy practices easier to understand and compare. Her role is part of a broader effort, not a claim of sole authorship. (CMU CUPS; CMU profile)

A 2024 paper co-authored by Cranor, Tianshi Li, Yuvraj Agarwal and Jason Hong describes Matcha, an Android Studio plugin for Google Play labels. It looks for first-party code paths through API calls and keywords, helps developers annotate data access and transmission, and uses an editable XML file to capture third-party SDK practices. The tool then produces a CSV that can be imported into Play Console. The important design choice is that the suggested evidence appears inside the developer’s coding environment, where someone can inspect, correct or reject it. (Matcha paper; open paper; project page)

Better accuracy has a cost

In a study of 12 developers creating labels for their own apps, 11 produced more accurate labels with Matcha than with the Play Console form. The participants worked across six Google Play apps. The Matcha task averaged 30 minutes; the console task averaged 9.8 minutes. That is a material trade-off: a more grounded process demanded about three times as much time in this study.

The result is promising evidence about workflow design, not proof of universal accuracy. The authors note that doing the console task first may have helped participants on the second task, their error ground truth was incomplete, and the sample may not represent large-company teams. Code analysis also cannot establish every server-side retention or later use after data leaves a device. Developers must still supply context and review the output.

A store label is therefore a mirror of what its rules expect a developer to know. Matcha tries to make that expectation more workable by creating a path from code and SDK evidence to a declaration. It does not turn the declaration into an audit. The durable contribution is a better place to notice uncertainty and correct a mismatch before it reaches the store page.