Case studies · Method

How the case studies are made.

Public data, and math. This brief is written for a reader who wants to check the work: what the sources are, what was counted and how, what the model is and is not, and where the limits lie. Everything in the case studies can be traced to a line here.

1. Purpose and scope.

The case studies describe the registered charitable sector of a community, or of a national body, from the public record, and use that description to explain what Public Benefit Canada does. They are descriptive. They make no causal claim, they predict nothing about any organization, and they are not used to select or approach organizations. The unit of analysis is the registered charity as it appears on its annual return; the population is charities that filed a return for the 2024 fiscal period and could be scored (79,678 of roughly 86,000 registered). Organizations that have stopped filing, which are the ones most likely to lose registration, are outside the population by construction, and every count in the case studies is therefore a floor.

2. Sources.

Canada Revenue Agency, List of charities and T3010 Registered Charity Information Returns, 2011 to 2024. Published annually on open.canada.ca under the Open Government Licence. The 2024 release used here was last revised by the CRA on 28 May 2026. Tables used: identification (name, business number, designation, category, mailing postal code), financial sections and Schedule 6 (revenue by source, expenditures, assets, liabilities, cash), directors and like officials (count, arm's-length share, dates), Schedule 3 (compensation and staffing), qualified donees (gifts to other charities), programs, and web presence. Filings are self-reported; ratios are winsorized at the 1st and 99th percentiles within each year.

Canada Gazette, Part I, notices of revocation of registration, 2011 to 5 September 2026. 13,639 unique business numbers, each typed by the notice's stated reason: failure to file, voluntary (at the charity's request), or following audit. Collected by a scraper from the Gazette's quarterly indexes and weekly issues; the file was last refreshed on 5 September 2026 and the last issue carrying a charity notice was 8 August 2026.

Elections Canada, 2023 representation order, postal code to electoral district. A charity is assigned to a federal riding by its mailing postal code. 321 of 343 ridings are covered; 22 are not yet in the crosswalk, and a case study that touches one of them says so in it. A mailing address is not a service area: a charity headquartered in one town may serve a region, and a national body's head office is in one riding.

Communities are defined as sets of ridings, named in each case study. National bodies are defined either by a shared business number (the Salvation Army: every registration under the Governing Council's number, which is exact) or by a pattern on the legal name (the United Church, the Anglican Church), with the pattern, its exclusions, and its known misses stated in the case study.

3. What is counted, and how.

Every figure in a community or body case study is produced by one script, case-study.py, from the scored 2024 file and the Gazette file. It counts: charities scored; charities by revenue band (six bands from under $25,000 to over $2 million, on total revenue reported); by charitable purpose (the CRA's category head: religion, education, relief of poverty, benefit to the community); boards by size, and the share with three members or fewer (the count of directors on the return); median revenue; the number and rate flagged by the model; and, where the Gazette has been matched to the community, revocations in the window stated, by reason. For national bodies it also counts revocations by year, 2011 to 2026, and, where the body has registrations held under a shared number, the same figures for held and independent registrations side by side.

Suppression. Any count under five is not printed; the case study says "fewer than five." This follows the practice of Statistics Canada for small cells and exists so that no organization can be identified by subtraction. It applies to totals as well as parts: a flag count or a revocation count under five is withheld, and its rate with it. And where a total is printed and any one of its parts is under five, every part is withheld, so that the small part cannot be recovered from the total. It applies to the stronger flag in most communities, and it is why some case studies give a count for one riding and not another, and why a small sector's case study has few numbers at all: in a sector of a few dozen charities, a count is a name. In the smallest sectors a revenue band under five is folded into its neighbour so that the ladder still prints, the share of boards with three members or fewer is withheld when the count is under five, and the board map is not drawn where the people on it would be fewer than five (added 8 September 2026).

Rounding. Rates are given to one decimal place in the case study and in the aggregate file; medians are rounded to the nearest thousand in prose only.

The riding-level output of the script for every case study is published as case-studies-aggregate.csv, so that the arithmetic can be checked without the organization-level file.

4. The model.

What it is. Canary v3 / Distress v3, built 1 July 2026, is a screening model for registered charities. It is a logistic regression with balanced class weights on 28 features derived from a charity's T3010 return (reserve ratio, months of cash, liabilities to assets, surplus margin, one-year revenue change, government and donation dependency, revenue concentration, program and compensation ratios, board size, arm's-length share and turnover, staffing, size, number of donees and programs, negative net assets, a small-board indicator, paid staff, web presence, foreign activity, foundation designation, age since reference, and three interaction terms for funding fragility, thin donor reliance, and structural deficit). It is trained on a panel of 431,032 charity-year observations across six snapshot years, 2011 to 2021.

Outcome. The outcome is all-cause revocation of registration published in the Gazette within three years of the snapshot. It is not distress and not closure. Failure-to-file revocations can lag an organization's actual death by years; some voluntary revocations are healthy amalgamations. The model finds organizations headed for revocation, which overlaps with but does not equal financial distress. This construct gap is the model's first limitation and is stated first on purpose.

Training and validation. Performance is reported fully out of time with strict separation of roles: train on 2011 to 2017 with recency weighting (six-year half-life), fit calibration on 2019, and evaluate discrimination, calibration, and flag characteristics on the 2021 cohort (43,261 charities), which neither stage had seen. The production model then retrains on 2011 to 2019 and calibrates on 2021, so that no year is used twice and no evaluation touches training data.

Calibration and flags. Raw scores are mapped to probabilities by Platt scaling fitted separately within each revenue band, so that the small-charity majority does not set the probability scale for large charities. A Canary flag requires both a relative and an absolute condition: calibrated risk at or above the 90th percentile within the band, and at least 1.25 times the band's observed base rate. The stronger flag (Distress) requires a revenue band of $75,000 or above, the 95th percentile, and at least twice the band base rate.

Performance, and what it means. On the 2021 cohort: area under the ROC curve 0.683 overall, and 0.51 to 0.65 within revenue bands. The three-year base rate was 1.39 per cent. The Canary flag had precision 2.5 per cent, recall 15 per cent, and lift 1.78; the stronger flag had precision 3.3 per cent and lift 2.34. Mean calibrated prediction was 1.22 per cent against an actual rate of 1.39, a calibration ratio of 0.88. In the 2024 scoring universe, mean calibrated risk is 1.40 per cent overall and 2.85 per cent among the 7,258 flagged.

In plain terms: the model ranks organizations better than chance, consistently, on years it never saw, and most of that ranking power comes from size. A flag roughly doubles the odds of revocation within three years, from about 1.4 in a hundred to about 3 in a hundred. Ninety-seven flagged organizations in a hundred are not revoked in the window. A flag is a statement that an organization's filings resemble the filings of organizations that later lost registration; it is not a prediction about that organization and it is not evidence of anything about its board.

Review. The method was tested by an external reviewer working from a data-free methodology export. The reviewer independently confirmed the within-band discrimination near chance and the out-of-time flag precision of 2.5 to 3.3 per cent. On that finding, organization-level prediction was retired as an objective of the work on 24 July 2026. The reviewer's report is available on request.

Reproducibility. All inputs are public. The complete pipeline is a single script with four stages that regenerates every number in the method paper from the raw panel; the model card records coefficients, scaling, per-band calibration parameters, and validation metrics. The method paper (Public Benefit Canada, working paper WP-2026-01, A Shrinking Commons), the model card, and the organization-level scored file are available to a reviewer under a data agreement whose only terms are that organization-level scores are not republished and no organization is approached on the strength of one.

5. The board networks.

The "Who sits on the boards" figure in each case study is built by a second script, case-network.py, from the directors table of the 2024 return. A person is matched across boards by exact normalized name; the figure keeps only charities that share at least one director with another, and only people on two or more of those boards. Every name, business number, and city is stripped before the file reaches the site: a charity keeps its purpose, revenue band, and riding or region; a person keeps only a count of boards.

Exact-name matching over-joins common names and under-joins spelling variants. Interlock counts are therefore estimates, and two director tables in our pipeline built by different dedup rules give counts that differ by roughly a sixth for the same region. We are moving to publishing a range; until then the case study gives the count from the file the reader can see drawn.

6. Limitations, in order of weight.

The outcome is revocation, not distress. The population is filers only, so every count is a floor. Within-band discrimination is modest; the flags concentrate risk about twofold and no more. A mailing address is not a service area, and 22 ridings are not yet in the postal crosswalk. Name-based selection of national bodies misses registrations whose legal names lack the pattern and is corrected by hand for false matches; the case study states the known misses. Director matching is by exact name. T3010 data are self-reported. Calibration is fitted on a single held-out year and would need refitting after a regime change such as another enforcement moratorium; the macro anchor is refreshed as each three-year Gazette window closes. The prose in each case study interprets the counts and, where it describes how a body's polity works, draws on general knowledge that is marked as such and is open to correction.

7. Responsible use.

Three rules govern, by a ruling of 24 July 2026. Organization-level scores and flags are never published and never shared as an assessment of any organization; only aggregates appear, with the model's version, date, population, and limitations. No flag generates outreach, a task, or a sales action; a conversation with an organization begins with that organization's own account, and only when it approaches us. Client-confidential information never enters the model. In the case studies a further rule applies: no organization is named unless its ending is complete and on the public record and the telling honours the people who carried it; a registration revoked for failure to file is named only where the organization's own decision to end was already on the public record, since the Gazette's reason column records the paperwork and not the decision (ruled 8 September 2026); otherwise it is never named.