A Test Collection for Sensitivity-Aware Search in Personal Information

This article discusses the creation of a test collection designed for sensitivity-aware search, focusing on the balance between information retrieval and personal privacy.

4 min readResearch

In the realm of information retrieval, traditional methods prioritize returning relevant documents based on user queries. However, many document collections also harbor sensitive personal data, raising concerns about privacy. To address this issue, there is a growing focus on Sensitivity-Aware Search (SAS) models that aim to deliver relevant results while safeguarding sensitive information. A key resource for developing these systems is a test collection that includes both sensitive and non-sensitive data, along with a variety of queries and assessments of relevance. The Enron email dataset serves as a practical example, containing genuine business emails, some of which include sensitive content. Yet, the original dataset lacks query formulations and relevance assessments. To remedy this, researchers have crowdsourced 150 queries across 50 topics, along with 11,471 relevance assessments for a carefully selected subset of Enron emails that have been tagged for sensitivity. Additionally, the collection has been enhanced using large language models (LLMs) to provide further assessments and sensitivity labels. Baseline performance metrics for relevance, sensitivity classification, and sensitivity-aware search have been established. The complete collection is accessible, including through the ir_datasets package, and pre-built indices are available on Huggingface for streamlined experimentation.

Research