Data Sources and Data Linkage
Research Methods in Health Sciences
Learning objectives for this lesson:
- Distinguish primary data from secondary data and describe the roles of data custodians, data stewards, data platforms and researchers.
- Describe the data available from Statistics Canada, including the confidential files held in Research Data Centres, and from the Canadian Institute for Health Information.
- Describe Population Data BC, its counterparts in other provinces such as ICES in Ontario, and the role of Health Data Research Network Canada in research across provinces.
- Outline the stages and contents of a data access request and plan realistically for its timelines and costs.
- Carry out a simple deterministic and probabilistic record linkage and interpret the decisions it produces.
- Explain missed matches and false matches and how uneven linkage error can bias the findings of a study.
- Explain de-identification techniques, the separation principle and the Five Safes framework, and describe how secure research environments and output checking protect privacy.
- Write a section on data sources, linkage and safeguards for a study’s data management plan, including a Five Safes table.
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University. It is the applied research methods course of the Public Health Assessment and Analysis series.
The Canadian Research Data Environment
Learning Objectives for this section
- Distinguish primary data collection from the secondary use of existing data, and describe the roles of data custodians, data stewards and data platforms.
- Describe the kinds of data Statistics Canada makes available and the three levels at which researchers can reach them: published tables, public use microdata files and confidential files in Research Data Centres.
- Describe the role of the Canadian Institute for Health Information and the main databases it holds on hospital stays, emergency department visits, continuing care and prescriptions.
- Describe Population Data BC and its counterparts in other provinces, including ICES in Ontario, and explain why research across provinces relies on networks such as Health Data Research Network Canada.
- Match a research question to the Canadian data source most likely to answer it.
Introduction
Lesson 8 showed how researchers collect quantitative data directly and introduced the main kinds of administrative health records: physician billing claims, hospital discharge abstracts, prescription records and vital statistics. This lesson asks where existing data are held in Canada, how a researcher obtains permission to use them, how records from different sources are joined, and what safeguards protect the people whose information is used.
Primary data are data that a research team collects for its own study, such as the answers to a survey it designed. Secondary data were collected for another purpose and are reused for research. A hospital records a diagnosis so that the patient can be treated and the hospital funded; a researcher who later counts hospital stays for heart failure is making secondary use of that record. Secondary data offer large numbers of people, long follow-up and information about people who never answer surveys, with no extra burden on the people described. Their limitations follow from their origin: they contain only what the original purpose required, their coding reflects administrative rules, and access takes time.
The Cedar Valley Social Connection Study is a fictional mixed-methods study used throughout this course. A team led by Dr. Maya Hart at a British Columbia university is working with the fictional Cedar Valley Health Authority, which serves about 210,000 residents, about 46,000 of them aged 65 and older. The team’s survey of older adults received 1,600 completed responses, and 392 respondents (24.5 percent) scored 6 or higher on the three-item UCLA Loneliness Scale. One question asks whether loneliness is associated with emergency department visits. The survey measures loneliness, but the visits are recorded in health system data, so the team must find where those records are held, apply for access, link them to the responses of people who consented, and protect the linked file. Each section of this lesson follows one of those steps.
1.1 Who Holds Health Data in Canada
Provinces and territories deliver most health care, so most records of physician visits, hospital stays, prescriptions and insurance registrations are created by provincial and territorial systems and held by ministries of health, health authorities and agencies. National organizations hold two other kinds of data. Statistics Canada, the national statistical agency, conducts the census and a large program of surveys. The Canadian Institute for Health Information (CIHI) receives records from the provinces and territories and organizes them to common national standards so that they can be compared.
Several roles recur in every access process, although organizations name them differently. A data custodian is the organization legally responsible for holding and protecting a data set, such as a ministry of health. A data steward is the person or body that decides, on the custodian’s behalf, whether a request to use the data should be approved. A data platform or data centre prepares data for research, links records from different sources and provides a secure place to analyze them. The researcher requests access, analyzes the data and is accountable for using them as approved. In British Columbia, the Ministry of Health is the custodian of physician billing records, data stewards review requests to use them, and Population Data BC has long been one of the platforms through which researchers request and analyze them.
1.2 Statistics Canada Surveys and Research Data Centres
Statistics Canada collects information under the Statistics Act, which requires it to keep information about individual people confidential. Its surveys use probability samples (Lesson 7) and weights to describe provinces and, for larger surveys, health regions.
The CCHS is a large cross-sectional survey of health status, health care use and the determinants of health among people aged 12 and older living in private households. It excludes some groups, including people living on First Nations reserves, people living in institutions such as long-term care homes, and full-time members of the Canadian Armed Forces. It asks about sense of belonging to the local community, which makes it useful in social connection research, but because it excludes long-term care residents it leaves out many of the oldest and frailest adults.
The CHMS combines an interview with direct physical measures, such as blood pressure and blood and urine samples, which provide measured values where other surveys rely on self-report, at the cost of a much smaller sample.
The census is conducted every five years, with a longer questionnaire for a sample of households (about one in four in recent censuses). Its tables describe the living arrangements of small areas, including the number of older adults who live alone.
The General Social Survey has covered a different theme in each cycle, such as caregiving, families and social identity, and in recent years Statistics Canada has asked about loneliness in some of its social surveys.
Three levels of access
| Level | What it contains | How researchers obtain it |
|---|---|---|
| Published tables | Counts, percentages and estimates already calculated by Statistics Canada | Free on the agency’s website |
| Public use microdata files (PUMFs) | One record per respondent, with identifying details removed or coarsened (for example, age in groups and broad geography) | Through university libraries under the Data Liberation Initiative, a partnership between Statistics Canada and post-secondary institutions |
| Confidential master files | Full record-level files with detailed variables and fine geography, and some linked files | Only for approved projects, inside a Research Data Centre or an approved virtual environment |
A public use microdata file describes real respondents, but identifying details have been removed or coarsened so that no one can be recognized. Students and researchers can usually analyze a PUMF obtained through a university library, although a file that reports only the province cannot answer a question about one health region.
Research Data Centres and the CRDCN
A research data centre (RDC) is a secure facility where approved researchers analyze confidential Statistics Canada files. The Canadian Research Data Centre Network (CRDCN), a partnership between Statistics Canada and Canadian universities, operates RDCs on more than thirty campuses, including campuses in British Columbia. RDCs hold survey master files, some administrative files such as tax and hospitalization records, and files from the agency’s Social Data Linkage Environment, which links survey, census and administrative records. Statistics Canada has, for example, linked responses from some health surveys to hospital and mortality records for respondents who agreed to linkage.
The researcher submits a proposal explaining why the confidential data are needed. Approved researchers undergo security screening and are sworn in as deemed employees of Statistics Canada, which places them under the confidentiality obligations of the Statistics Act. They work on computers that cannot send files out, and every result they wish to take away is reviewed by Statistics Canada staff to confirm that it cannot reveal information about an individual. The network has introduced a virtual RDC for remote work under similar controls, and fees depend on the institution and type of project, so students should ask local RDC staff about current arrangements.
What this means for Cedar Valley
The Cedar Valley team could use census tables to describe how many older adults in the region live alone and the CCHS to compare community belonging among older adults in British Columbia with other provinces. Neither can answer the team’s main question, because Statistics Canada files cannot be joined to the Cedar Valley survey.
1.3 The Canadian Institute for Health Information
CIHI is an independent, not-for-profit organization, created in 1994, that collects health system data from the provinces, territories and health care organizations and publishes information about Canada’s health systems. Its value to researchers lies in standardization. Hospitals code diagnoses with the Canadian version of the International Classification of Diseases (ICD-10-CA) and procedures with the Canadian Classification of Health Interventions (CCI), so a hospital stay for pneumonia is recorded the same way in Nova Scotia and British Columbia.
CIHI’s public reports and interactive tools give aggregate results, such as hospitalization rates by province, without any application. Researchers who need custom tables or record-level files submit a data request, which CIHI reviews for privacy and which can involve cost-recovery fees. Record-level CIHI files are de-identified and are not usually linked to a researcher’s own survey. Studies that need that kind of linkage more often use the provincial version of the same records through a provincial data centre, because the province holds the health insurance numbers that make linkage possible.
1.4 Population Data BC and Its Counterparts in Other Provinces
Population Data BC (PopData) is a data and education resource hosted at the University of British Columbia. It has long helped researchers request, link and analyze British Columbia administrative data. These include the Ministry of Health, whose data sets cover physician services billed to the Medical Services Plan, hospital discharges, prescriptions recorded in PharmaNet and registration files, along with sources such as vital statistics, the cancer registry and education records. Researchers submit a data access request through PopData’s Data Access Unit, data stewards review it, and approved researchers analyze the data remotely in a secure environment. PopData can also link data that researchers collected themselves, such as survey responses, to administrative records, provided that participants explicitly agreed to the linkage on a consent form that meets the data providers’ requirements and has been approved by a research ethics board.
Arrangements in British Columbia have been changing. In 2024 the Ministry of Health began accepting academic data requests through the Health Data Platform BC, and since August 2025 new research projects that request Ministry of Health data have been handled under that program through a shared PopData and Health Data Platform BC request process. Analysis is moving from PopData’s Secure Research Environment to a cloud-based Trusted Analysis Environment. Students should check the current process before planning a project; the principles in this lesson apply in either arrangement.
| Organization | Province | What it does |
|---|---|---|
| Population Data BC, with the Health Data Platform BC | British Columbia | Supports requests for and linkage of provincial administrative data, including researcher-collected data with consent |
| ICES (originally the Institute for Clinical Evaluative Sciences) | Ontario | An independent, not-for-profit research institute that holds linked, coded health data for Ontario residents under the province’s health privacy law; researchers outside ICES reach the data through its Data and Analytic Services |
| Manitoba Centre for Health Policy (MCHP) | Manitoba | A research centre at the University of Manitoba that maintains the Manitoba Population Research Data Repository |
| Health Data Nova Scotia | Nova Scotia | A data centre at Dalhousie University that supports access to linked provincial health data |
Other provinces and territories have their own arrangements, which also change over time.
Research across provinces
Provincial privacy laws and data agreements usually require record-level health data to stay in the province, so a study comparing provinces cannot simply pool their files. One solution is distributed analysis: the team writes a common protocol and analysis code, analysts in each province run the code on their own data, and only summary results are combined. The Canadian Network for Observational Drug Effect Studies (CNODES) has used this approach to study the safety of medications across several provinces. Health Data Research Network Canada (HDRN Canada) is a non-profit network of provincial, territorial and pan-Canadian data organizations that works to make research across regions easier. Its Data Access Support Hub offers researchers a single entry point and a common request form for data from more than one region, an inventory of available data sets, and an inventory of algorithms (tested definitions of conditions written in terms of administrative codes).
1.5 Cohort Studies, Registries and Indigenous-Governed Data
Some sources were created for research from the start. The Canadian Longitudinal Study on Aging (CLSA) follows more than 50,000 adults who were aged 45 to 85 when recruited and collects information on health, function and social participation over many years; researchers apply through its own access process. Disease registries, such as provincial cancer registries, record every diagnosed case of a condition in a defined population.
Data about First Nations, Inuit and Métis people raise questions of governance that go beyond privacy law. The First Nations Information Governance Centre conducts the First Nations Regional Health Survey and holds it under the OCAP® principles of ownership, control, access and possession described in Lesson 4. Provincial records also include First Nations people, and the use of data that identify them can require approval under agreements with First Nations governance bodies. For the Cedar Valley team, the fictional Cedar Valley First Nations Health Centre takes part in decisions about whether and how data about its members are requested, analyzed and reported.
For each Cedar Valley question, name the source you would try first and one limitation of that source. (1) How many adults aged 65 and older in the region live alone? (2) Are respondents who score 6 or higher on the loneliness scale more likely than others to visit an emergency department in the following year? (3) How does the rate of hospital stays among older adults in British Columbia compare with other provinces? (4) Is social participation in mid-life associated with later changes in health among Canadian adults?
(1) Census tables, which count people who live alone; living alone is a different concept from loneliness. (2) The team’s survey linked to provincial records through Population Data BC; only respondents who consented can be included, and emergency department coverage must be checked for the years needed. (3) CIHI reports or a CIHI data request using the Discharge Abstract Database; differences between provinces can reflect how services are organized as well as differences in health. (4) The Canadian Longitudinal Study on Aging, which follows the same adults over time; its volunteers may be healthier than the population as a whole.
Section 2 turns from where the data are held to how a researcher obtains permission to use them.
Reflection
A health authority planner asks you which data source to use for each of three questions. (a) What proportion of adults aged 65 and older in each community of the region live alone? (b) Do older adults in British Columbia who report a weak sense of belonging to their local community have more hospital stays than those who report a strong sense of belonging? (c) How do emergency department visit rates among older adults in British Columbia compare with those in Ontario? Four sources are available. Census tables give counts of living arrangements for small areas and are free online. The Canadian Community Health Survey asks people aged 12 and older living in private households about their sense of community belonging; it excludes people living in long-term care, and its confidential files, including some files linked to hospital records for respondents who agreed to linkage, are available only in Research Data Centres after a proposal is approved. CIHI’s National Ambulatory Care Reporting System holds emergency department visits in a standard national format, but coverage of emergency departments differs between provinces. Population Data BC supports access to linked British Columbia administrative data, which contain no measure of community belonging. For each question, choose a source, explain why it fits, and name one limitation you would report.
(a) Census tables are the best source, because they count living arrangements for small areas, including each community, and they are free. One limitation is that living alone is a different concept from loneliness, so the tables describe a risk factor and say nothing about how people feel. Counts for very small communities may also be rounded.
(b) The Canadian Community Health Survey linked to hospital records, analyzed in a Research Data Centre, fits because it is the only source that combines a measure of community belonging with hospital stays. Population Data BC holds the hospital records but no belonging measure. Limitations include the exclusion of people in long-term care, who are among the heaviest users of hospitals, and the restriction to respondents who agreed to linkage. Access requires an approved proposal, so the work would take months.
(c) CIHI’s National Ambulatory Care Reporting System fits because it records visits in the same format in both provinces. Before comparing, I would check that emergency departments in both provinces report for the years needed, because coverage differs. I would also note that differences in rates may reflect how emergency and primary care are organized in each province as well as differences in health.
Minimum 20 characters required.
Question 1: A researcher counts hospital stays for heart failure using discharge abstracts that hospitals created for patient care and funding. Which statement describes these data?
Question 2: A researcher needs detailed geography to study loneliness among older adults in a single health region using Statistics Canada survey data. Which level of access does the project require?
Question 3: Why do studies that link a researcher’s own survey to health records in British Columbia usually work through a provincial data platform instead of CIHI?
Question 4: A team wants to compare hospital use among older adults in British Columbia, Ontario and Manitoba, and each province requires its record-level data to stay in the province. Which approach fits?
Applying for Access: Data Access Requests, Timelines and Costs
Learning Objectives for this section
- Describe the typical stages of a data access request, from the first enquiry to the closure of the project.
- Prepare the main parts of a data access request, including a study population definition and a justified list of data sets and variables.
- Explain who reviews a request and on what grounds, including data stewards, research ethics boards and Indigenous governance bodies.
- Explain why access to administrative data takes months, identify the costs a project should budget for, and build data access into a project timeline.
Introduction
Administrative health data cannot be downloaded. A researcher who wants to use them must make a formal application, usually called a data access request, which describes the research, the people and records it needs, and the safeguards that will protect them. The request is reviewed by the people responsible for the data, and access is granted only for the approved purpose. This section describes that process in general terms. Each organization has its own forms and rules, which change over time, so the details of any real application come from the organization’s current guidance and its data access staff.
Planning for access begins well before the request is submitted. The Cedar Valley team (a fictional study) knew from the start that it wanted to link survey responses to health records. It therefore wrote the linkage into its consent form (Lesson 5), asked respondents for their personal health number in the survey (Lesson 8), and contacted Population Data BC while the survey was still being designed. Of the 1,600 adults aged 65 and older who completed the survey, 1,312 (82.0 percent) consented to linkage and provided a personal health number. Those 1,312 people are the population for which the team can request linked records.
2.1 The Stages of a Data Access Request
Processes differ between organizations, but most follow a similar sequence. Figure 2.1 shows eight typical stages. Population Data BC, for example, describes its process in stages that run from planning the request through data steward review, contracts and privacy training, data release, analysis and pre-publication review, to project closure.
In the enquiry and intake stage, the researcher contacts the organization’s data access staff, describes the question and learns which data sets might answer it. In the feasibility and cost stage, the researcher confirms that the data exist for the needed years and population, reads the data set documentation (often called metadata or a data dictionary), and asks for a cost estimate, which funders may want to see in a grant application. The ethics stage obtains approval from a research ethics board (REB). The submission stage completes the request form. In data steward review, the people responsible for each data set decide whether to approve the request, often after asking questions. The agreements and training stage covers the signed research agreement, confidentiality undertakings from every team member, privacy training and the creation of accounts. In the linkage and release stage, the platform links and extracts the records and places them in a secure environment. The final stage covers analysis, output review and closure: results are checked before they leave the environment, some organizations review publications before release, and at the end of the project the data are destroyed or retained as the agreement specifies.
2.2 What a Data Access Request Contains
A data access request is a short research proposal written for a particular audience. Data stewards want to know whether the use is appropriate, whether the amount of data requested is the minimum the question needs, and whether the people and settings involved will keep the data safe. The tabs below describe the usual parts of a request, with the Cedar Valley version of each.
What to write. State the research question, the reason it matters for the health of the population, and how the findings will be used. Reviewers look for a clear public benefit and a question that the requested data can answer.
Cedar Valley. The team asks whether adults aged 65 and older who report loneliness have more emergency department visits in the year after the survey than those who do not, after accounting for age, sex, prior health conditions and community. The findings will inform the health authority’s planning of social connection programs.
What to write. Define exactly who is included, using criteria that the data provider can apply. This definition is often called the cohort definition. A vague definition, such as “older adults in the region”, cannot be turned into an extract.
Cedar Valley. The population is the 1,312 survey respondents who consented to linkage and provided a personal health number. No other residents are requested.
What to write. List each data set, the years required and each variable, with a one-line justification for every variable. Request the least detailed version that answers the question, for example age in years rather than full date of birth.
Cedar Valley. The worked example below lists the team’s data sets and variables.
What to write. Explain which files will be linked, what identifiers will be used and who will handle them, and state whether participants consented to linkage or whether a waiver of consent is requested.
Cedar Valley. Respondents gave written consent to linkage. The team will send the identifiers of consenting respondents to the platform’s linkage staff, separately from their survey answers.
What to write. Name every person who will see the data and describe their role and training. Describe where the analysis will take place.
Cedar Valley. Dr. Hart and the graduate research assistant will analyze the data inside the secure environment. The community research associate and the advisory group will see only approved summary results.
What to write. Summarize the planned analysis and the kinds of results that will be released, and describe how findings will be shared with the communities involved.
Cedar Valley. The team will compare the proportion of lonely and non-lonely respondents with an emergency department visit and will report results by community only where counts are large enough to protect privacy.
Data minimization
The principle that runs through every part of a request is data minimization: a project should receive only the people, years and variables that its question requires, at the least identifying level of detail that will work. Population Data BC’s guidance puts the idea directly: the selection of data fields in an extract must be restrictive. Minimization reduces the harm that would follow a breach, and it makes approval easier, because every extra variable is another item a data steward must be persuaded to release.
The table lists the data sets the team requested for the 1,312 consenting respondents. The years are expressed relative to each person’s survey date.
| Data set | Period | Main variables | Justification |
|---|---|---|---|
| Registration and demographic file | Two years before to one year after | Age in years, sex, health service area of residence, months of coverage | Confirms residence and coverage during follow-up and supplies adjustment variables |
| Physician billing records (Medical Services Plan) | Two years before to one year after | Service date, diagnostic code, type of practitioner | Identifies prior chronic conditions and primary care contact before the survey |
| Hospital discharge abstracts | Two years before to one year after | Admission and discharge dates, main diagnosis | Identifies prior hospital stays, a marker of poorer health |
| Emergency department visit records | One year after | Visit date, triage level, main diagnosis | Measures the outcome |
| Death records (vital statistics) | One year after | Date of death | Ends follow-up for people who died |
The team also recorded what it chose not to request. It did not ask for full dates of birth, because age in years was sufficient, or for full postal codes, because the health service area identified each community. It did not ask for physicians’ identities, which the question does not require. Before submitting, the team confirmed with data access staff that emergency department records were available for every community in the region for the years needed.
2.3 Who Reviews a Request, and on What Grounds
Several groups review a request, and each looks at it from a different angle.
Data stewards decide whether the use is permitted under the law and agreements that govern each data set and whether it is appropriate. In British Columbia, the Freedom of Information and Protection of Privacy Act sets the conditions under which public bodies may disclose personal information for research, and Ontario’s Personal Health Information Protection Act plays a similar role for ICES. Stewards consider the public benefit of the research, whether the data requested are the minimum needed, whether consent was obtained or its absence is justified, and whether the people and settings involved are secure.
Research ethics boards review the study under the Tri-Council Policy Statement (TCPS 2), whose rules on privacy, secondary use of information and consent Lesson 5 described. TCPS 2 contains a specific article on data linkage (Article 5.7). It requires REB approval before linkage takes place, asks researchers to describe the data to be linked and the likelihood that linkage will create identifiable information, and, where it will, asks researchers to show that the linkage is essential to the research and that appropriate security measures will protect the information. Data providers usually require proof of current ethics approval before they release data, and any change to the request later may require an amendment to both approvals.
Indigenous governance bodies have authority over data about their communities under agreements and under the principles described in Lesson 4. In the Cedar Valley study, the fictional Cedar Valley First Nations Health Centre reviewed the request, agreed with the team on how results about First Nations participants would be reported, and asked to review any such results before release.
The researcher’s institution signs the research agreement, because the university, as well as the researcher, takes on legal obligations for the data.
2.4 Timelines and Costs
Access to linked administrative data takes time. Population Data BC describes the request, approval and data provisioning process as one that can take some months, and requests involving several data providers, new linkages or researcher-collected data tend to take longer. Researchers should treat any study that depends on linked administrative data as a long-term undertaking, and a study that must finish within a few months is rarely feasible unless the data are already available to the team.
Requests are often delayed for predictable reasons. The study population may be defined in terms that the data provider cannot apply. The list of variables may be long and poorly justified, which prompts questions from stewards. Ethics approval may be missing, expired or written for a different version of the study. A request may involve several data providers, each with its own review. Agreements may wait for signatures from the university. Data preparation and linkage are done in a queue with other projects. Changes made after approval, such as adding a variable, usually require an amendment.
A researcher can shorten the wait by contacting data access staff early, attending an intake meeting with a clear question and population, reading the data documentation before choosing variables, keeping the request as small as the question allows, obtaining ethics approval that explicitly covers the linkage, and starting the request while primary data collection is still under way.
Access also costs money. Data platforms commonly charge fees to recover the cost of preparing, linking and extracting data and of providing a secure analysis environment, and the size of those fees depends on the complexity of the request. Statistics Canada RDC access can involve fees that depend on the institution and project, and CIHI custom requests can involve cost-recovery charges. A project budget should also include the time of an analyst who can work with large administrative files, the time all team members need for privacy training, and the cost of amendments. A cost estimate obtained during the feasibility stage allows these items to be included in a grant application.
| Activity | Cedar Valley planning allowance |
|---|---|
| Intake meeting, feasibility check and cost estimate | Months 1 to 2, while the survey is being built |
| Ethics approval covering linkage and submission of the request | Months 2 to 4 |
| Data steward review, questions and revisions | Months 4 to 7 |
| Agreements, confidentiality undertakings, privacy training and accounts | Months 7 to 9 |
| Linkage, data preparation and release | Months 10 to 12 |
| Analysis and output review | From month 13 |
These allowances are the fictional team’s own planning assumptions. Actual times depend on the request and the organization, and the team confirmed its assumptions with data access staff at the intake meeting. The months are counted from the intake meeting, which the team held in month 3 of the project, so access approval falls in project month 11 and the linked data arrive in month 14, as the Gantt chart in Lesson 6 shows. Lesson 6 also showed how to place such allowances in a project timeline so that analysis and reporting are scheduled realistically.
Suppose you want to study whether older adults who live alone are less likely to fill their prescriptions after a hospital stay. Write one sentence of justification for each of three variables you would request (for example, the dispensing date from prescription records). For each one, state the least detailed version that would still answer the question. Then name one variable you would deliberately leave out and explain why.
Once access is approved, the platform must find each consenting respondent in the administrative files. Section 3 describes how that record linkage is done and what happens when it goes wrong.
Reflection
Two research assistants draft a request for linked administrative data for a small pilot study. They plan to link a survey of 400 family caregivers to provincial physician billing records. Their draft states that the study population is “caregivers in the region”; requests full dates of birth, full postal codes, every physician billing record from 2000 to the present, and the names of the physicians each person saw; does not mention consent to linkage; notes that ethics approval has not yet been sought; says the analysis will be done on one assistant’s personal laptop; and expects the data within six weeks so that the pilot can be finished within four months. A data access request usually describes its purpose, study population, data and variables, linkage and consent, people and security, and analysis and outputs, and it passes through enquiry, feasibility and cost, ethics approval, submission, data steward review, agreements and training, linkage and release, and output review. Identify at least four problems with the draft and propose a revised plan.
The study population is too vague for a data provider to apply. It should be defined as the caregivers among the 400 survey respondents who consented to linkage and supplied a personal health number. The variable list breaks the principle of data minimization: age in years and a broad area of residence would replace full dates of birth and postal codes, the years should be limited to the period the question needs (for example, two years before and one year after the survey), and physicians’ names should be dropped because the question does not need them. Consent is missing: linkage of researcher-collected data to provincial health records requires each participant’s explicit agreement to the linkage on a consent form that meets the data providers’ requirements and has REB approval. Ethics approval must come first, because TCPS 2 requires REB approval before linkage and providers ask for proof of approval. A personal laptop is unacceptable; the data would stay in the platform’s secure environment. Finally, six weeks is unrealistic, because requests, approvals and data release commonly take months.
A revised plan would keep the pilot feasible by preparing the request for later submission while the team analyzes a public use microdata file now. If the linkage goes ahead, the team would contact data access staff first, obtain a cost estimate, and schedule the linked analysis for a later phase of the study.
Minimum 20 characters required.
Question 1: What is the main purpose of the feasibility stage of a data access request?
Question 2: Which change to a draft data request best applies the principle of data minimization?
Question 3: What does TCPS 2 Article 5.7 require of researchers who propose to link data?
Question 4: A research assistant hopes to obtain linked provincial administrative data and finish the analysis within three months. What is the most accurate advice?
Record Linkage: Deterministic and Probabilistic Methods and Linkage Error
Learning Objectives for this section
- Define record linkage and describe the identifiers used to link health records and the errors they commonly contain.
- Carry out a simple deterministic linkage using an exact rule and explain how stepwise rules extend it.
- Explain how probabilistic linkage adds agreement and disagreement weights into a total weight, and apply upper and lower thresholds with clerical review.
- Distinguish missed matches from false matches and explain how each can bias the results of a study.
- Describe what researchers can do to detect, reduce and report linkage error.
Introduction
After its request was approved, the Cedar Valley team (a fictional study) needed the platform to find each of its 1,312 consenting respondents in the provincial files. That task is record linkage: bringing together records that belong to the same person (or family, address or organization) from different sources, or from within one source. Linkage allows a survey answer about loneliness to sit beside a record of an emergency department visit, and it is also a source of error that is invisible in the final data set.
3.1 What Record Linkage Is
Halbert Dunn (1946), an American public health statistician, used the term “record linkage” to describe assembling the records of a person’s life, from birth to death, into a single “book of life”. Howard Newcombe and colleagues (1959), working in Canada, showed that a computer could link birth and marriage records automatically by weighing how strongly agreement on names and other details suggested that two records belonged together. Ivan Fellegi and Alan Sunter (1969), statisticians at what is now Statistics Canada, gave probabilistic linkage the mathematical form it still has. Fellegi later served as Chief Statistician of Canada.
A linkage key is the identifier or combination of identifiers used to decide whether two records belong together, such as a health number, or a surname combined with a date of birth. A match is a pair of records that truly belong to the same person. A link is the decision, made by a rule or a reviewer, to treat two records as belonging to the same person. Linkage error is the gap between them.
3.2 Identifiers and Why They Disagree
Health records in British Columbia carry a personal health number (PHN), a lifetime identifier assigned to each person registered for provincial health insurance. It is the strongest linkage key available, because each person has one and no two people should share one. Names, birth dates, sex and postal codes can all disagree between two records of the same person.
| Identifier | Common reasons two records of the same person disagree |
|---|---|
| Personal health number | Digits mistyped or transposed when copied into a survey or form; number missing |
| Surname | Spelling variants (MacLeod and McLeod); change of name on marriage; hyphenated names recorded in part; different transliterations from other alphabets |
| Given name | Short forms and nicknames (William and Bill); use of a middle name; names recorded in a different order |
| Date of birth | Day and month reversed; typing errors; missing values replaced by a default date |
| Sex | Recording errors; changes in recorded sex or gender over time |
| Postal code | Moves between the dates of the two records; a mailing address in one record and a home address in the other |
These errors are unevenly spread. People who move often, whose names are often misspelled or transliterated, or who have changed their names are more likely to have records that disagree. Older adults who have recently moved to a smaller town, the group at the centre of one Cedar Valley qualitative question, are one example.
3.3 Deterministic Linkage
Deterministic linkage links two records when they agree exactly on a specified linkage key or set of identifiers. The simplest rule uses one high-quality identifier: link two records if their personal health numbers are identical. A stepwise (or multi-pass) deterministic linkage applies a sequence of rules, starting with the strictest; records left unlinked by the first rule are tried against a second, such as agreement on surname, given name, date of birth and sex.
Each row pairs a survey record with a hospital record, showing the last four digits of each health number. The final column shows the truth, which the linkage team would not know.
| Pair | Survey record | Hospital record | PHN (survey / hospital) | Truth |
|---|---|---|---|---|
| 1 | Helen Ostrowski, F, 14 Mar 1951, V0X 2B1 | Helen Ostrowski, F, 14 Mar 1951, V0X 2B1 | 4471 / 4471 | Same person |
| 2 | Donald MacLeod, M, 2 Jun 1948, V0X 1C5 | Donald McLeod, M, 2 Jun 1948, V0X 1C5 | 2093 / 2093 | Same person |
| 3 | Margaret Chen, F, 30 Sep 1955, V0X 3K2 | Margaret Chen, F, 30 Sep 1955, V0X 3K2 | 8812 / 8821 | Same person |
| 4 | William Thomas, M, 11 Jan 1944, V0X 1C5 | Bill Thomas, M, 11 Jan 1944, V1A 4R9 | 5530 / 5530 | Same person |
| 5 | Ruth Okafor, F, 22 Jul 1950, V0X 2B1 | Joan Okafor, F, 22 Jul 1950, V0X 2B1 | 6604 / 6648 | Different people (twin sisters) |
| 6 | Helen Ostrowski, F, 14 Mar 1951, V0X 2B1 | Peter Ostrowski, M, 9 Nov 1949, V0X 2B1 | 4471 / 3307 | Different people (spouses) |
Rule: link if the personal health numbers are identical. Pairs 1, 2 and 4 link, despite the spelling difference in pair 2 and the different given name and postal code in pair 4, because the rule looks only at the health number. Pairs 5 and 6 are correctly left unlinked. Pair 3 is a missed match: the respondent transposed two digits when writing her health number. A second pass requiring agreement on surname, given name, date of birth and sex would recover pair 3.
Deterministic linkage is transparent and works very well when a unique identifier is recorded accurately in both files. A single error in the linkage key prevents a link, however, because a strict rule cannot tell one transposed digit from a different person.
3.4 Probabilistic Linkage
Probabilistic linkage compares several identifiers and asks how strongly the pattern of agreement and disagreement supports the conclusion that two records belong to the same person. It is used when no unique identifier is available in both files, or when the identifier is incomplete. Following Fellegi and Sunter (1969), each identifier is described by two probabilities.
The two probabilities behind each weight
The m-probability is the probability that an identifier agrees when the two records truly belong to the same person; it is a little below 1 because of recording errors, nicknames and moves. The u-probability is the probability that the identifier agrees by chance when the records belong to different people; for sex it is about one half, and for a full date of birth among older adults it is very small.
An agreement weight is large when m is high and u is low. Software converts the ratio m ÷ u to a logarithmic scale so that weights can be added; each additional point means the observed pattern is twice as likely for a true match as for a non-match. A disagreement weight, built from (1 − m) ÷ (1 − u), is negative.
Total weight = the sum of the agreement or disagreement weight for each identifier.
Software usually estimates the probabilities from the data and applies blocking, comparing only pairs that agree on a simple variable such as year of birth, over several passes.
Suppose the hospital file had no health numbers, so the team must link on name, date of birth, sex and postal code. The first table shows illustrative probabilities and weights; the second scores the six pairs.
| Identifier | m | u | Weight if it agrees | Weight if it disagrees |
|---|---|---|---|---|
| Surname | 0.95 | 0.01 | +6.6 | −4.3 |
| Given name | 0.90 | 0.01 | +6.5 | −3.3 |
| Date of birth | 0.97 | 0.0005 | +10.9 | −5.1 |
| Sex | 0.99 | 0.50 | +1.0 | −5.6 |
| Postal code | 0.80 | 0.01 | +6.3 | −2.3 |
| Pair | Surname | Given name | Birth date | Sex | Postal code | Total | Decision | Truth |
|---|---|---|---|---|---|---|---|---|
| 1 | +6.6 | +6.5 | +10.9 | +1.0 | +6.3 | 31.3 | Link | Same |
| 2 | −4.3 | +6.5 | +10.9 | +1.0 | +6.3 | 20.4 | Link | Same |
| 3 | +6.6 | +6.5 | +10.9 | +1.0 | +6.3 | 31.3 | Link | Same |
| 4 | +6.6 | −3.3 | +10.9 | +1.0 | −2.3 | 12.9 | Clerical review | Same |
| 5 | +6.6 | −3.3 | +10.9 | +1.0 | +6.3 | 21.5 | Link | Different |
| 6 | +6.6 | −3.3 | −5.1 | −5.6 | +6.3 | −1.1 | Non-link | Different |
The team set an upper threshold of 15 (pairs at or above it are linked) and a lower threshold of 5 (pairs below it are not linked). Pairs in between go to clerical review, in which a trained person examines the records and decides. Agreement on a full birth date earns far more weight than agreement on sex, because different people rarely share a birth date and often share a sex. Pair 3 now links, because the transposed health number plays no part. The reviewer accepts pair 4, because Bill is a common short form of William and the postal code shows a move within the region. Pair 5 is a false match: twin sisters who live together agree on everything except their given names, and their total of 21.5 passes the upper threshold.
Moving the thresholds trades one error for another. Raising the upper threshold to 25 would send pairs 2 and 5 to clerical review, where the reviewer would probably reject the twins, at the cost of more manual review; lowering the thresholds would reduce missed matches and increase false matches.
3.5 Linkage Error and How It Is Measured
A missed match (a false negative) occurs when two records of the same person are not linked. A false match (a false positive) occurs when records of two different people are linked. When the truth is known for a sample of pairs, for example from a careful manual check, two measures summarize performance. Sensitivity is the proportion of true matches that were linked. Positive predictive value is the proportion of links that are true matches.
| Method in the worked example | True matches linked | Links that were true matches | Errors |
|---|---|---|---|
| Deterministic, health number only | 3 of 4 (sensitivity 75 percent) | 3 of 3 (positive predictive value 100 percent) | One missed match (pair 3) |
| Probabilistic, with clerical review | 4 of 4 (sensitivity 100 percent) | 4 of 5 (positive predictive value 80 percent) | One false match (pair 5) |
Six pairs are too few to judge a method, but they show the usual pattern: deterministic rules on a strong identifier tend to produce few false matches and some missed matches, while probabilistic methods recover more true matches at the risk of some false ones. Many linkage units combine the two, with a deterministic pass on the health number followed by probabilistic passes for the records that remain.
3.6 The Consequences of Linkage Error
The consequences of linkage error depend on how unlinked records are handled and on whether errors are spread evenly across the groups being compared (Harron et al., 2017; Doidge & Harron, 2019).
Suppose that 90 of the 298 consenting lonely respondents (30.2 percent) and 200 of the 1,014 other consenting respondents (19.7 percent) truly had at least one emergency department visit in the following year, a gap of 10.5 percentage points. Now suppose a weaker linkage that relies on postal codes misses 10 percent of lonely respondents’ visits, because they moved more often, and 3 percent of other respondents’ visits. The analysis would count 81 lonely respondents with a visit (90 minus 9) and 194 others (200 minus 6). The observed proportions would be 81 ÷ 298 = 27.2 percent and 194 ÷ 1,014 = 19.1 percent, a gap of about 8 percentage points. The association would appear about one quarter smaller, entirely because of the linkage.
In the actual (fictional) Cedar Valley linkage, the platform linked 1,268 of the 1,312 consenting respondents in a deterministic pass on health number and date of birth and 30 more in probabilistic passes; 5 of 8 pairs sent to clerical review were accepted. In total, 1,303 respondents (99.3 percent) were linked. Five of the 9 unlinked respondents were lonely (1.7 percent of 298) and 4 were not (0.4 percent of 1,014), a small error in the same direction as the example, which the team reported.
3.7 What Researchers Can Do
Researchers usually receive linked data without identifiers, so they depend on the linkage unit for information about quality. The GUILD guidance (Gilbert et al., 2018) asks those who link data to share their methods, the proportion of records linked and how linkage rates differ between groups.
Analysts can compare linked and unlinked people on every variable in the original file, such as age, community and loneliness score. A difference warns that missed matches may bias the results.
Events after a recorded death can signal false matches. A sensitivity analysis repeats the main analysis under different assumptions, for example using only the most certain links; a conclusion that holds under each is less likely to be an artefact of the linkage.
The RECORD statement (Benchimol et al., 2015), an extension of the STROBE reporting guideline for studies using routinely collected health data, asks authors to describe the linkage methods and the quality of the linkage. Lesson 12 introduces reporting guidelines more fully.
Use these weights (agreement: surname +6.6, given name +6.5, date of birth +10.9, sex +1.0, postal code +6.3; disagreement: surname −4.3, given name −3.3, date of birth −5.1, sex −5.6, postal code −2.3) and thresholds of 5 and 15. Pair A agrees on surname, given name, sex and postal code, but its birth dates are 04 Jul 1949 and 07 Apr 1949. Pair B agrees on given name, date of birth and sex, and disagrees on surname and postal code. Calculate each total, state the decision, and say what a clerical reviewer would look for.
Pair A: 6.6 + 6.5 − 5.1 + 1.0 + 6.3 = 15.3, just above the upper threshold, so the pair is linked. The reversed day and month is a common recording error, so a link is plausible, although a cautious team might review pairs this close to the threshold. Pair B: −4.3 + 6.5 + 10.9 + 1.0 − 2.3 = 11.8, which goes to clerical review. The reviewer would check for a spelling variant or change of surname and for a move within the region.
Section 4 describes how a linked file, which reveals more than any of its sources, is protected.
Reflection
A linkage uses these weights. Agreement: surname +6.6, given name +6.5, date of birth +10.9, sex +1.0, postal code +6.3. Disagreement: surname −4.3, given name −3.3, date of birth −5.1, sex −5.6, postal code −2.3. Pairs scoring 15 or more are linked, pairs scoring below 5 are not linked, and pairs in between go to clerical review. Pair X agrees on surname, given name, date of birth and sex and disagrees on postal code. Pair Y agrees on surname, sex and postal code and disagrees on given name and date of birth. Pair Z agrees on date of birth and sex and disagrees on surname, given name and postal code. (1) Calculate each total and state the decision. (2) In the full study, which compares hospital visits between older adults who moved in the past year and those who did not, 12 percent of movers’ records and 2 percent of other participants’ records fail to link, and unlinked participants are counted as having no visits. Explain the likely effect on the comparison and describe two things the researchers could do about it.
(1) Pair X: 6.6 + 6.5 + 10.9 + 1.0 − 2.3 = 22.7, which is above 15, so it is linked; a changed postal code is consistent with a move. Pair Y: 6.6 − 3.3 − 5.1 + 1.0 + 6.3 = 5.5, which falls just inside the review zone. A reviewer would notice that the two records share a surname and an address but have different given names and birth dates, which suggests two members of one household, and would probably reject the link. Pair Z: −4.3 − 3.3 + 10.9 + 1.0 − 2.3 = 2.0, which is below 5, so it is not linked.
(2) Movers’ visits would be undercounted far more than other participants’ visits, because 12 percent of movers would appear to have no visits whatever happened to them. If movers truly have more visits, the observed difference would shrink and could disappear; if they have fewer, it would be exaggerated. The bias comes from the linkage, so a reader could not detect it from the results. The researchers could ask the linkage unit for linkage rates by mover status and compare linked and unlinked participants on the survey variables. They could also repeat the analysis using only participants whose records linked, or add a deterministic pass on health numbers if those are available, and report the linkage methods and rates following the RECORD statement.
Minimum 20 characters required.
Question 1: In record linkage, what is the difference between a match and a link?
Question 2: Why does agreement on full date of birth receive a much larger weight than agreement on sex?
Question 3: Using a lower threshold of 5 and an upper threshold of 15, a pair of records scores 12.9. What happens to it?
Question 4: In a study, 10 percent of lonely respondents’ records and 3 percent of other respondents’ records fail to link, and unlinked respondents are counted as having no emergency department visits. What is the likely effect?
Privacy Safeguards: De-identification, the Five Safes and Secure Research Environments
Learning Objectives for this section
- Distinguish direct identifiers from indirect identifiers and explain why combinations of indirect identifiers can identify people, especially in small communities.
- Describe the main de-identification techniques, including removal, pseudonymization, generalization, suppression and aggregation, and explain the separation principle used in linkage.
- Explain the Five Safes framework and use it to assess a data access arrangement.
- Describe the features of a secure research environment and the checks applied to results before they are released.
- Apply these safeguards in the data management plan for a study.
Introduction
A linked file can reveal more about a person than any of its sources. The Cedar Valley survey (a fictional study) records how lonely a respondent feels; the health records show when that person went to an emergency department and why. Joined together, they form a detailed account of a person’s life that the person shared on the understanding that it would be protected. Lesson 5 introduced privacy and confidentiality under TCPS 2, the categories of identifiable information, and data management plans. This section describes the specific safeguards that make the research use of linked administrative data possible: reducing how identifiable the data are, controlling who uses them and where, and checking every result before it is released.
4.1 How People Are Identified
A direct identifier identifies a person on its own: a name, a personal health number, a street address, a telephone number or an email address. An indirect identifier, also called a quasi-identifier, does not identify anyone on its own but can do so in combination with other information. Date of birth, postal code, sex, ethnicity, occupation, a rare diagnosis and the dates of hospital stays are all indirect identifiers.
The power of combinations is easy to underestimate. Latanya Sweeney (2000), a computer scientist, estimated that about 87 percent of the population of the United States could be uniquely identified by three items: five-digit ZIP code, sex and full date of birth. Removing names from a file therefore does little to protect people if the file keeps detailed dates and locations. TCPS 2 reflects this by distinguishing directly identifying, indirectly identifying, coded, anonymized and anonymous information (Lesson 5), and by treating information as identifiable whenever it could reasonably be expected to identify a person, alone or in combination with other available information.
Risk is higher in small populations. In a city of 90,000 people, a 91-year-old woman who visited an emergency department in March is one of many. In a small community such as the fictional Kestrel Lake in the Cedar Valley region, a neighbour who knows those three facts might recognize her. This is why rural communities, small First Nations communities and people with rare conditions need particular care in the release of results, and why the release rules described below often require small groups to be combined.
4.2 De-identification Techniques
De-identification is the set of techniques that reduce the chance that a person can be identified from a data set, while keeping the data useful for the research question. It reduces risk to a low level without removing it entirely, so it is always combined with the other safeguards in this section. Khaled El Emam (2013), a Canadian researcher in this field, describes de-identification as a process of measuring the risk of re-identification and applying techniques until the risk is acceptably low for the setting in which the data will be used.
| Technique | What it does | Cedar Valley example |
|---|---|---|
| Removal | Deletes direct identifiers from the analysis file | Names, health numbers and street addresses never enter the analysis file |
| Pseudonymization (coding) | Replaces identifiers with a code; the key that connects codes to people is held elsewhere | Each respondent carries a project-specific study number that means nothing outside this project |
| Generalization | Replaces precise values with broader categories | Age in years replaces date of birth; the health service area replaces the postal code |
| Suppression | Removes or groups rare values | Ages above 95 are recorded as “95 or older” |
| Aggregation | Releases counts or summaries for groups instead of records | The advisory group sees tables of counts, never individual records |
| Rounding and perturbation | Alters released numbers slightly so that small counts cannot be read exactly | Statistics Canada, for example, randomly rounds many census counts |
The separation principle
Linkage needs identifiers, and analysis needs content, but no single person needs both. The separation principle organizes linkage around that fact. Data providers send identifying information (names, health numbers, birth dates) to a linkage unit, which uses them to work out which records belong together and assigns each person a project-specific study number. The content (loneliness scores, visit dates, diagnoses) travels separately, without names or numbers, and is joined using the study numbers inside a secure environment. Linkage staff therefore see identifiers without health content, and researchers see health content without identifiers. Because the study numbers are specific to one project, files released to different projects cannot be joined to each other.
4.3 The Five Safes Framework
De-identification protects the data themselves. Other safeguards control the project, the people, the setting and the results. The Five Safes framework, developed by Felix Ritchie and colleagues from work on secure data access at the United Kingdom’s Office for National Statistics (Desai et al., 2016), organizes all of these safeguards as five questions. Statistical agencies and data centres in several countries use it to design and explain their access arrangements, and it is a practical tool for planning any study that uses sensitive data. Click each card to read the question it asks.
The five dimensions work as a set. When one is strong, others can be lighter, and when one must be weak, others must compensate. A public use microdata file has very safe data, because detail has been removed, so students and researchers at participating institutions can analyze it on an ordinary computer. A confidential master file in a Research Data Centre has much less safe data, because it keeps detail that researchers need, so access requires screened people, a controlled setting and checked outputs. The table compares three arrangements described in Section 1.
| Dimension | Public use microdata file | Research Data Centre | Provincial linked data in a secure environment |
|---|---|---|---|
| Projects | Any lawful use under the licence terms | Proposal approved before access | Data access request approved by data stewards, with ethics approval |
| People | Users at participating institutions | Screened researchers sworn in as deemed employees | Named team members with privacy training and confidentiality undertakings |
| Settings | The user’s own computer | Secure physical or virtual centre | Remote secure research environment |
| Data | Detail removed or coarsened | Detailed record-level files | De-identified, linked records with study numbers |
| Outputs | No check required | Every output reviewed by Statistics Canada staff | Outputs checked against the platform’s rules |
The Five Safes is a planning framework. Privacy law, data agreements, TCPS 2 and Indigenous data governance under principles such as OCAP® apply alongside it, and a First Nations partner may set conditions on projects, people and outputs that go further than any of the five questions require.
4.4 Secure Research Environments and Output Checking
A secure research environment (also called a trusted research environment) is a controlled computing system in which approved researchers analyze sensitive data without being able to remove them. Population Data BC’s Secure Research Environment, the cloud-based Trusted Analysis Environment that is replacing it, the analysis environments at ICES and the virtual Research Data Centre all follow this model, although their details differ. Common features include remote access to a virtual desktop rather than a copy of the data, strong authentication when logging in, no internet, email or copy-and-paste out of the environment, a record of what each user does, and a single controlled route by which files enter and results leave.
That single route is where output checking happens. Before a table, figure or model result leaves the environment, it is checked for disclosure risk, either by the researcher against the platform’s rules, by platform staff, or both. The main risks are listed below.
A table cell based on very few people can identify them, especially when combined with other information. Many organizations require counts below a set threshold (for example, below five) to be suppressed or combined with other categories. The exact rule differs between organizations, so researchers use the rule that applies to their data.
Two tables that are each safe can be unsafe together. If one table reports 40 people and a second, otherwise identical table excludes one small community and reports 38, subtraction reveals a count of 2. Checking outputs as a set prevents this.
A maximum or minimum value often belongs to a single person. Scatter plots and lists of outliers can show individual records. These are usually replaced with percentiles or summary measures.
Regression results are usually low risk, but a model fitted to a very small subgroup, or one that includes a category with very few people, can reveal information about those people.
The graduate research assistant prepared the following table of linked respondents with three or more emergency department visits in the year after the survey. The team’s rule, confirmed with the platform, is that no released cell may be between 1 and 4.
| Group | Cedar City | Riverside | North Bench | Kestrel Lake | Total |
|---|---|---|---|---|---|
| Lonely (score 6 or higher) | 9 | 4 | 3 | 2 | 18 |
| Not lonely | 12 | 5 | 4 | 1 | 22 |
Five cells fall between 1 and 4, so the table cannot be released. Suppressing those five cells alone would leave a problem, because the totals and the remaining cells would allow some suppressed values to be worked out by subtraction. The team instead combined the three smaller communities.
| Group | Cedar City | Other communities | Total |
|---|---|---|---|
| Lonely (score 6 or higher) | 9 | 9 | 18 |
| Not lonely | 12 | 10 | 22 |
Every cell is now 5 or more. The team also checked that no other output in the same report showed Kestrel Lake separately in a way that would allow differencing, and it sent results that concerned First Nations participants to the Cedar Valley First Nations Health Centre for review before release, as the partners had agreed.
4.5 Worked Example: The Cedar Valley Five Safes Plan
The Cedar Valley team summarized its safeguards in a Five Safes table that it attached to its data management plan and its data access request. A plan of this kind shows reviewers, partners and participants how each risk is handled.
| Dimension | Cedar Valley safeguards (fictional) |
|---|---|
| Safe projects | The research ethics board approved the study, including linkage; the data stewards approved the data access request; the Cedar Valley First Nations Health Centre agreed to the use of data about its members; the purpose is to inform the health authority’s social connection programs. |
| Safe people | Only Dr. Hart and the graduate research assistant can open the linked data; both completed the platform’s privacy training and signed confidentiality undertakings; the university signed the research agreement; the community research associate and the advisory group see only released results. |
| Safe settings | All analysis of linked data takes place in the platform’s secure environment; the team stores the survey identifiers on the university’s secure server in a separate location from the survey answers, as its data management plan describes. |
| Safe data | Linked records carry project-specific study numbers; age in years replaces date of birth; the health service area replaces the postal code; ages above 95 are grouped. |
| Safe outputs | No released cell may be between 1 and 4; small communities are combined; the platform checks outputs; results about First Nations participants are reviewed by the health centre before release. |
Writing the data sources, linkage and safeguards section
A data management plan can include a section titled “Data sources, linkage and safeguards” with three parts. First, it states whether the study uses any existing data source; if it does, it names the source, the organization that holds it and the route by which access is requested, using the descriptions in Sections 1 and 2. Second, it states whether the study links any data; if it does, it names the identifiers, explains who holds them, and confirms that the consent form asks participants for permission to link. Third, it includes a five-row Five Safes table modelled on the Cedar Valley example in Section 4.5. When a study uses only data the team collects itself, the table still applies: it describes how identifiers are separated from responses, where the data are stored, who sees them, and the rule for small numbers in reported results.
Reflection
A researcher has linked survey and hospital data for 250 older adults living in five communities, the smallest of which has about 300 residents. She plans two things. First, she will email a spreadsheet of the linked records to the six members of a community advisory group so that they can look for patterns; the spreadsheet contains a study number, full date of birth, full postal code, sex, community, diagnosis codes and the date of each emergency department visit. The advisory group members have not completed privacy training or signed confidentiality undertakings. Second, she will publish a table of emergency department visits by community and five-year age group in which several cells contain 1, 2 or 3 people. The Five Safes framework asks whether the project, the people, the setting, the data and the outputs are safe. Identify a problem under each of the five dimensions and propose a revised plan.
Safe projects: sharing record-level data with the advisory group is probably outside the use approved by the data stewards and the ethics board, which named who would see the data. Safe people: the advisory group members have no training and no confidentiality undertakings, so they should not receive record-level data at all. Safe settings: email and personal computers offer no control over copying or onward sharing; linked data should stay inside the secure research environment. Safe data: full date of birth, full postal code and visit dates would allow neighbours in a community of 300 people to recognize individuals, even without names. Safe outputs: cells of 1 to 3 people could identify individuals, especially in the smallest community.
In a revised plan, the researcher would keep all record-level data in the secure environment and analyze it there herself. She would generalize ages to broader groups and combine the smaller communities, so that every cell in the published table met the platform’s threshold, and she would check that no other table allowed suppressed values to be recovered by subtraction. She would then meet the advisory group to interpret the checked, aggregate tables together, which keeps their knowledge in the analysis without exposing anyone’s records.
Minimum 20 characters required.
Question 1: Which of the following is an indirect identifier, also called a quasi-identifier?
Question 2: Under the separation principle used in data linkage, who sees the names and health numbers used to link the records?
Question 3: A public use microdata file can be analyzed on an ordinary computer, while confidential master files require a Research Data Centre. How does the Five Safes framework explain the difference?
Question 4: A table shows 2 lonely respondents in Kestrel Lake with three or more emergency department visits. Under a rule that no released cell may be between 1 and 4, what is the best response?
Final Assessment
Bringing It All Together
This lesson followed the path that existing data take from the systems that create them to a research result. In Canada, most administrative health records are created by provincial and territorial health systems, while Statistics Canada holds the census and national surveys and CIHI organizes records from the provinces to common national standards. Researchers reach these data through published tables, public use microdata files, Research Data Centres, CIHI data requests and provincial platforms such as Population Data BC and ICES, and networks such as HDRN Canada support research that crosses provincial borders.
Access depends on a data access request that defines the study population precisely, justifies each variable, explains the linkage and consent, names the people involved and describes the secure setting. Data stewards, research ethics boards and Indigenous governance bodies each review the request, and the process commonly takes months and carries costs that belong in the project budget. Once approved, records are joined by deterministic rules, probabilistic weights or both. Both methods make errors, and missed matches and false matches can bias a study’s results when they fall unevenly on the groups being compared.
The linked file is protected by layers of safeguards. De-identification and the separation principle reduce what any one person can see, and the Five Safes framework asks whether the project, the people, the setting, the data and the outputs are each safe enough. Secure research environments and output checking keep record-level data inside a controlled system and ensure that released results cannot identify anyone, which matters most in small communities such as those in the fictional Cedar Valley region.
Key Takeaways from this lesson
- Secondary data were collected for another purpose, such as patient care or billing, and offer large populations and long follow-up at the cost of limited variables and slow access.
- Provinces and territories hold most administrative health records, Statistics Canada holds census and survey data, and CIHI organizes provincial records to common national standards.
- Statistics Canada data are available as published tables, as public use microdata files and, for approved projects, as confidential master files in Research Data Centres.
- Population Data BC, ICES, the Manitoba Centre for Health Policy and Health Data Nova Scotia support access to linked provincial data, and HDRN Canada and distributed analysis support research across provinces.
- A data access request defines the study population exactly, justifies each variable under the principle of data minimization, and describes linkage, consent, people, setting and outputs.
- Data stewards, research ethics boards and Indigenous governance bodies review requests on different grounds, and TCPS 2 Article 5.7 requires ethics approval before data linkage.
- Access to linked administrative data commonly takes months and carries fees, so it belongs in the project timeline and budget from the start.
- Deterministic linkage applies exact rules to identifiers, while probabilistic linkage adds agreement and disagreement weights and uses thresholds with clerical review.
- Missed matches undercount events and false matches attach the wrong events, and either can bias comparisons when it falls unevenly on the groups being compared.
- De-identification, the separation principle, the Five Safes framework, secure research environments and output checking work together to protect the people whose records are used.
Core Concepts Reviewed
Section 1: primary and secondary data, data custodians and stewards, Statistics Canada surveys, public use microdata files, Research Data Centres and the CRDCN, CIHI databases, Population Data BC and the Health Data Platform BC, ICES and other provincial centres, HDRN Canada and distributed analysis.
Section 2: the stages of a data access request, the study population definition, data minimization, data steward and ethics review including TCPS 2 Article 5.7, Indigenous data governance, and timelines and costs.
Section 3: record linkage, linkage keys, matches and links, deterministic and stepwise linkage, m- and u-probabilities, agreement weights, thresholds and clerical review, missed and false matches, sensitivity and positive predictive value, and linkage bias.
Section 4: direct and indirect identifiers, de-identification techniques, the separation principle, the Five Safes framework, secure research environments, output checking and differencing.
The final reflection asks you to plan the data sources, access, linkage and safeguards for a new study from start to finish.
Reflection
A community organization in British Columbia runs a weekly meal program for older adults. It keeps attendance records with each participant’s name, date of birth, sex and postal code, but no health numbers. Of 600 participants, 480 signed a consent form allowing their attendance records to be linked to provincial hospital records for research; some participants attend with a spouse who shares their surname and address. A researcher wants to know whether regular attendance is associated with fewer hospital stays in the following year. Write a short data plan (about 200 words) that (1) names who holds each kind of data and the general route to access, (2) lists three things the data access request must contain, (3) states which linkage approach the identifiers allow and one linkage error to expect, with its likely effect, and (4) gives one safeguard for each of the Five Safes (projects, people, settings, data and outputs).
(1) The community organization holds the attendance records, and the Ministry of Health is the custodian of the hospital discharge records, which the researcher would request through the shared Population Data BC and Health Data Platform BC process, after an intake meeting and a feasibility and cost check. (2) The request must define the population as the 480 consenting participants, list each variable with a justification (admission and discharge dates, main diagnosis, age in years and health service area), and include evidence of written consent and ethics approval that explicitly covers linkage, as TCPS 2 Article 5.7 requires.
(3) Without health numbers, the linkage must be probabilistic, or stepwise deterministic, on name, date of birth, sex and postal code. Spouses who attend together share a surname and postal code, so a false match between them is possible; it would attach one spouse’s hospital stays to the other and blur any real difference between regular and occasional attenders. Missed matches among people who moved are also likely and would undercount their stays.
(4) Projects: ethics and data steward approval with a clear public benefit. People: only named, trained analysts who have signed confidentiality undertakings. Settings: analysis only in the secure research environment. Data: study numbers, age in years and health service area replace identifiers. Outputs: small cells combined or suppressed and every table checked before release.
Minimum 30 characters required.
Final Knowledge Assessment
Question 1: The Cedar Valley team needs emergency department visits for survey respondents who consented to linkage. Which source and route fit best?
Question 2: Which statement about Statistics Canada Research Data Centres is accurate?
Question 3: A researcher wants to study emergency department visits using a CIHI database. Which database is it, and what should be checked first?
Question 4: What does Health Data Research Network Canada’s Data Access Support Hub provide?
Question 5: A data steward asks why a request includes full postal codes. Which reply applies data minimization?
Question 6: Which reviewers decide whether a use of data is permitted under the law and agreements governing each data set and whether the data requested are the minimum needed?
Question 7: In the lesson’s worked example, pair 3 had a transposed digit in the survey’s health number. The deterministic rule missed it, but probabilistic linkage linked it. Why?
Question 8: In the worked example, twin sisters who live together scored 21.5 and were linked. What kind of error is this, and which change would most likely catch it?
Question 9: A linkage correctly links all 4 true matches in a sample and makes 5 links in total. What are its sensitivity and positive predictive value?
Question 10: In the Cedar Valley survey, 76.0 percent of lonely respondents and 83.9 percent of other respondents consented to linkage. What does this imply for the linked file?
Question 11: Latanya Sweeney estimated that most of the United States population could be uniquely identified by sex combined with which two items?
Question 12: Which safeguard belongs to the safe settings dimension of the Five Safes?
Question 13: Why might suppressing only the small cells in a table still fail to protect privacy?
Question 14: Which statement about the Five Safes framework is accurate?
Question 15: A study uses only interviews that the research team conducts. How does this lesson apply to its data management plan?
Glossary: Key Terms, People & Frameworks
📚 Reference page, available throughout the lesson
These terms, organizations and people appear in this lesson on Canadian data sources, record linkage and privacy safeguards.