Skip to content

(Draft) Registered access tier for the AnVIL Data Explorer #4954

Description

@NoopDog

1. The problem

Most AnVIL datasets are controlled access. A researcher submits a data access request, waits weeks or months for approval, and only then can see who is in the study. Today the Explorer shows nothing at the donor, sample, or file level for datasets you do not have access to.

The dataset-level view helps a little. Each dataset carries rolled-up tags: if any donor has diabetes, the dataset is tagged "diabetes"; if any donor is female, it is tagged "female". So you can find datasets tagged both "diabetes" and "female". What you cannot find out is whether any dataset has donors who are both female and have diabetes, let alone how many.

Result: researchers apply for datasets that turn out not to have the cohort they need, or skip datasets that do. The application queue is slow and the cost of a wrong guess is high.

2. What we are building

A registered access tier. A researcher who signs in with a RAS-linked account (Researcher Auth Service, the NIH identity system now required for Terra) and accepts a click-through agreement gets one new capability: filters combine at the donor level across every dataset, including ones they cannot open, and they see roughly how many donors in each dataset match.

"Roughly" is the whole design. Counts are shown as ranges, never exact numbers, and anything under 50 is shown as "Under 50". Those two rules protect donors, and they are not negotiable in the UI.

The product is organized into two sections, one per user goal:

  1. Find studies and request access. The registered tier lives entirely here. Query, rank datasets by how many donors match the user's query, and link each controlled-access study to its dbGaP request page.
  2. My studies: build a cohort and export. The current catalog, scoped to the datasets the user has access to. Exact counts, entity tabs, dataset selection, and the existing Export flow.
Image

The existing three-state Access facet (Public, Required, Granted) drives every action in both sections.

3. Who it is for

  • Primary: a researcher planning a data access request who needs to know, before applying, that the datasets contain a cohort like "Hispanic female donors with diabetes and whole genome sequencing".
  • Secondary: program staff (NHGRI, AnVIL) who need the tier to be defensible to a data access committee (DAC, the group that approves requests).
  • Not for: the public tier. Logged-out behavior is unchanged apart from a prompt to sign in.

4. Goals

  1. A registered user can express a donor-level query using the existing filter facets and see per-dataset matching counts.
  2. The path from "logged out" to "searching at donor level" is two steps: sign in, accept.
  3. The user always understands which section they are in (datasets to request vs data they have access to) and why some counts are ranges.
  4. Every disclosure-controlled number is inside Find studies, is a range from a fixed ladder, and never distinguishes 0 from 1–49.

5. Disclosure Control Strawman

  • Datasets the user cannot open show counts only; no donor, sample, or file rows.
  • Filters are categorical. "Request access" opens the study's dbGaP page, as it does today.

5.1. The bin ladder

Labels as they appear on screen:

Under 50   50+   100+   250+   500+   1,000+   2,500+   5,000+   10,000+

Each label is the lower bound of its range and means "at least this many and fewer than the next step". A tooltip can spell out the full range ("1,000 to 2,499"). Use the same labels on every surface.
Developer notes: disclosure rules behind Find studies

These come from the disclosure-control design and are the reason the tier is defensible to a data access committee. They apply to Find studies only. Design should treat them as fixed; the visible consequences are already described in sections 7 and 8.

5.2. Key Requirements

Constraint What it means for the UI
Minimum count of 50 Any matching count under 50 is shown as "Under 50". Zero and 49 look identical. Do not use 0, a dash, or an empty cell for these.
Bins, not numbers Every matching count is a range from one ladder. No hover-to-reveal, no export of matching counts.
At most two diagnosis/phenotype concepts per query The query builder refuses a third concept. The refusal message is fixed text about the rule, never about the data ("this combination would be too small" is forbidden because it leaks). Both "any" and "all" modes count toward the cap.
Under-threshold datasets are not ranked When sorting by matching count, everything under 50 falls into one collapsed, unranked group. Ordering them would leak which are nearer 49. Ties within a bin sort alphabetically, never by the hidden count.
Same appearance for suppressed cells everywhere Totals and per-dataset counts use the same "Under 50" treatment.
Totals are derived, never computed The "matching" totals line is a function of the per-dataset bins already on screen: the dataset count is the number of rows at 50+, and the donor total is the exact sum of the displayed lower bounds (under-50 rows count as zero), shown as "at least N" and labeled as a sum of the ranges' lower bounds. A true cross-dataset total, binned or fuzzed, is a new cell and is not released. No biosample total until per-dataset biosample bins exist.
Facet values show dataset counts only In the sidebar and any Add-filter picker, a value shows the number of datasets tagged with it (public index). Never donor or sample counts, and never counts conditioned on the current query. TODO: review if we can show any counts on facets.
Dataset-level facets are only dataset properties Access, consent group, organism, identifier. Everything under DONOR, BIOSAMPLE, or FILE, including data modality and file format, describes a donor, so it is a donor-level filter: applying one adds the matching-donor range to each dataset row and sorts the table by it.
Never grey a value because of the data A greyed or refused facet value is greyed by a rule (the concept cap), never because the resulting count would be small.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions