Every caveat on the methods page, worked out: the reasoning, the robustness checks and the numbers behind each one.
Bounding the headline against how crashes get reported
The outcome behind every number on this site is a police-reported collision, and reporting practice shifted during the study. The severity mix drifts visibly across years: the fail-to-remain share ran 15.8 percent in 2019, 8.9 in 2022, and 20.8 in 2024, while the property-damage-only share swung between 65 and 84 percent. If a change in how crashes get reported, rather than in how many occur, drove the headline finding, that would be a real problem, so we bound it rather than assume it away.
The logic runs on the two ends of the reporting-sensitivity spectrum. An injury collision is close to immune to reporting practice, someone is hurt, police attend, no threshold applies. A property-damage-only collision is the opposite: exactly where thresholds and intake changes bite. If the within-site decline shows up at both extremes, reporting drift cannot be what produces it. It does.
| Strip | Median change (per site-month) | Sites improved | p |
|---|---|---|---|
| All collisions | −0.104 | 58.5% | 1.1 × 10⁻⁵ |
| Property-damage only | −0.088 | 58.8% | 8.4 × 10⁻⁵ |
| Injury (least reporting-sensitive) | −0.030 | 63.7% | 7.8 × 10⁻⁶ |
All three strips decline, measured as a per-site rate difference rather than a ratio, since a ratio has a floor at −100 percent that a rare event hits easily. The injury strip, the one least exposed to reporting drift, shows the most consistent direction of the three. A citywide reporting shift is also ruled out at the design level too: matched against pseudo-control locations on the same calendar clock, camera cells declined 3.9 percentage points more than their matches, a comparison detailed under regression to the mean in the Limits section below, and a reporting change common to the whole city would shift both groups alike. It cannot produce that gap.
▶ Verify the three strips in R, recomputed live from the published file.
Two qualifications. The symmetric zero-exclusion rule applied to crash counts elsewhere in this report does not hold for a rare-event strip like injuries: keeping only sites with a non-zero during-window count selects the windows where an injury happened to occur by chance, which flips the ratio positive on its own, a selection artefact we document but do not use. And the pooled aggregate rate ratios for these strips sit slightly above one, injuries at 1.007, property-damage at 1.048, which is the familiar placement-selection pattern from the formal tests above, not a contradiction: cameras went to sites and eras with rising crash counts, and a pooled aggregate inherits that selection where a within-site comparison removes it. The within-site figures are the ones that decline at both ends of the reporting spectrum, and they are the ones we report.
Limits we cannot argue away
We think readers are best served by the limits stated as plainly as the findings.
A single effect size. The estimate shifts with the research design, from roughly zero to −10 percent. Our judgment of −4 to −7 percent is a judgment, and we have tried to show the work rather than declare a number.
How much of the post-removal rise was camera-specific. The direction is clear; the magnitude is entangled with the post-pandemic recovery and cannot be separated from it with these data.
How much belongs to the camera alone. 383 of 520 sites sat within 250 metres of a radar sign, and at 193 of those, 37 percent of the network, the sign's deployment overlapped the camera's, so what we call the camera effect there is strictly a bundle. Splitting by baseline volume gives four cells. On the busier half of the network, where the crashes and the prevented crashes concentrate, cameras without a sign show −9.2 percent (p = 0.057) against −7.9 for bundled sites (p = 0.0006), statistically indistinguishable (group difference p = 0.51). On the low-traffic half the cells diverge, cameras alone at +3.2 percent against −12.1 for bundled sites, though both readings are noisy at that volume and neither difference clears significance on its own. The sign does not appear necessary on the half of the network carrying most of the benefit. No cross-site split can separate the constants that remain regardless: school-zone signage, markings and zone speed limits persist at nearly every site.
Direct severe-injury effects. KSI events near camera sites are too rare to test, so those figures are imputed from the injury effect.
Geocoding precision. A camera's coordinates mark the nearest street intersection, not the pole itself, which introduces some imprecision into a 250-metre buffer. The buffer-distance check on the methods page, which finds the effect present and significant from 100 metres to a kilometre, suggests this does not materially change the result.
Concurrent interventions. Toronto's Vision Zero programme, launched in 2016, brought road redesigns, protected cycling infrastructure and pedestrian signal changes that matured over 2018 to 2022, overlapping substantially with the camera rollout. We do not know which of the 520 sites also received a concurrent redesign, and where one did, our within-site estimate attributes the combined effect to the camera. The departure rebound argues against Vision Zero carrying the whole effect, since a road redesign would not disappear when a camera does, but the two interventions may also interact, so this remains an unresolved confound rather than a bounded one.
Regression to the mean. Cameras went to locations with a documented crash and speed problem, and an unusually bad period at any location tends to be followed by a more ordinary one regardless of what changes there. We built 2,637 pseudo-control locations, 250-metre grid cells with a 2014-to-2020 crash rate similar to a camera site's and at least 500 metres from any camera, and compared their decline against camera cells over the same calendar window. The untreated cells fell 17 percent; camera cells fell 15, nearly indistinguishable, which confirms that much of the raw decline is a citywide trend rather than a camera effect. Pairing each of 419 camera cells with its closest-matched control cell narrows that to a camera-specific difference of 3.9 percentage points (Wilcoxon p = 0.006), and that gap held roughly flat across five baseline-crash-rate quintiles, from −1.5 to +2.1 points, which argues against regression to the mean driving it: the effect would be expected to concentrate at the highest-crash quintile, and it does not. Pushing the pre-period baseline forward from 2014 to 2017 pulls the headline within-site estimate from −10.2 to −6.9 percent, still significant (p < 0.001, 302 of 520 sites improved) and closer to the matched-control floor. Taken together we read the defensible range as 4 to 7 percent, and the departure rebound remains the clearest single argument against regression to the mean as the whole explanation: it cannot account for crashes rising again once a camera leaves.
Risk per kilometre driven. Traffic-volume data for the period do not exist in public form, so every reduction in this report is a reduction in crash counts, and if enforcement pushed traffic elsewhere rather than slowing it, a count-based estimate would conflate the two. Three checks argue against that reading. Crashes rise again once a camera leaves, and diverted traffic would be expected to return with it, which is the opposite of what we see. Radar signs near 410 camera sites recorded a median 67,000 vehicles a month, heavily trafficked roads with no convenient detour. And low-volume sites showed a larger reduction than high-volume ones, −13 percent against −5, the reverse of the pattern diversion would produce, since a low-traffic street is easier to route around than a busy one. Matching 406 sites to a nearby radar sign and converting counts to a rate per 10,000 vehicles left the estimate essentially unchanged, −11.1 against −10.9 percent (Wilcoxon p = 7.5 × 10⁻⁷), and traffic volume itself showed no correlation with the size of the reduction (ρ = 0.085, p = 0.087).
That reporting practices held still. About 16 percent of collision records lack usable coordinates, ranging from 78.5 percent valid in 2019 to 89.0 in 2016, and the ones that geocode carry a higher injury rate, 14.3 percent against 10.4 for the ones that do not (χ² = 1,371.8). Every figure in this report necessarily uses only the geocoded share, so it modestly overstates the injury component. That would distort a within-site comparison only if the geocoding rate itself shifted around enforcement, and it does not: 84.0 percent before the programme, 82.7 during, 84.0 after. And Waterloo's 2024 data show signs of a coding change. Both are named and bounded; neither can be fully corrected.
Which sites will respond next time. Site-level outcomes are dominated by noise at the crash counts involved. Our best cross-validated model explains about 11 percent of the site-to-site variation.
Short enforcement periods. Many deployments ran only 3 to 6 months, which may be too brief for deterrence to fully develop, so this report may understate what a longer deployment achieves. Sites that returned for a second or third deployment did show larger reductions, which is consistent with that reading.
Speed data taken at one point in time. The Watch Your Speed measurements linked to camera sites are a single snapshot, not a before-and-after pair, so they show what speeds look like at a camera site rather than how much a camera changed them. The Howard and Rothman (2025) SickKids and TMU study measures the change directly and finds speeding down 45 percent.
Seasonal composition. Enforcement start dates cluster in July and November, and Toronto's crash rate varies by roughly 25 percent across the calendar year, lowest in April and highest in November, so before and during periods could in principle sample different seasons. A seasonally adjusted version of the within-site test shifted the estimate by 0.7 percentage points, from −10.3 to −9.6 percent. The concern is real; the effect on the finding is not.
Camera sites sit close together. The median distance between neighbouring cameras is 352 metres, and 93 percent of sites have another camera within a kilometre, so overlapping 250-metre buffers could in principle treat correlated observations as independent and understate the uncertainty. Grouping the network into 311 spatial clusters and re-running the paired test on cluster-level means gave a stronger result, a median −9.6 percent reduction against −10.2 percent on individual sites, and restricting to the 76 cameras at least 750 metres from any neighbour gave −20.1 percent. One open question remains: whether the not-yet-treated sites the Callaway–Sant'Anna estimator leans on are themselves free of spillover from a nearby active camera. We cannot verify that they are.
Multiple testing. This report runs roughly 35 to 40 hypothesis tests across several families. With that many, some nominally significant results are expected by chance alone. Applying a false-discovery-rate correction within each family, every core finding survives, the departure rebound, the paired test, the difference-in-differences estimate, the staggered design, and the strongest predictors. Two secondary results do not survive correction, transit commuting as a predictor and one wellbeing sub-score, and both were already presented as suggestive rather than conclusive.