In research circles, “open” has become a synonym for “good”. Open access, open source, open data. I share the instinct. But after years of building language datasets, I've learned that releasing data openly is a decision with consequences, not a default.
When openness exposes
Social media posts collected for research can be traced back to their authors with surprising ease, even after names are removed. For activists, journalists, or minority communities in the region, that's not an abstract privacy concern. It can be a matter of safety.
Anonymization is a promise we often can't keep. Consent is a promise we can.
Designing for consent
When we built the LAHJA corpus, every contributor chose to take part and can withdraw their text at any time. That made collection slower and more expensive. It also meant we could explain, honestly, how the data would be used.
- Tiered access: open for aggregate statistics, gated for raw text.
- Data statements documenting who is represented, and who isn't.
- A clear process for removal requests, with named people responsible.
What policymakers can do
Funders increasingly require open data. They should also fund the governance that makes openness responsible — the review boards, the access committees, the maintenance. Openness without stewardship isn't generosity. It's abdication.
