We're tracking and scoring maps throughout redistricting. Check out our grades at our Redistricting Report Card ↗️

Methodology

Producing a report card requires two steps. First, we calculate metrics for a plan using a standardized approach. Second, we use those metrics to assign grades according to a rubric that we apply to all states.

Metrics

Partisan Composition

We use recent statewide elections to calculate a partisan index of each voting precinct, which we can then use to calculate district-by-district estimates by two-party Democratic vote share (DVS). (Use of DVS is a party-agnostic measure, and is chosen as a convention; use of estimated Republican vote share throughout this calculation gives the same results.) The tallied elections are confined to the statewide offices of U.S. President, U.S. Senate, and Governor. Vote counts are averaged to determine voteshares. For some states, precinct data is not available, and the base unit is instead a Census block group.

Using the Democratic voteshare, we are able to calculate the expected Democratic seat share. We evaluate this in two ways: integer and real Democratic seats. Integer Democratic seats are calculated simply by counting the number of districts with DVS above 50%. Real Democratic seats are calculated by summing the Democratic win probability across each district. This accounts for uncertainty and ensures that closely divided districts contribute proportionally rather than being treated as guaranteed wins or losses. The win probability is calculated using a cumulative t-distribution: P(Democratic Win|DVS) = tCDF(DVS-0.5 / 0.037;3). The parameter 3.7% was selected to match the overall win probabilities of 2012-2020 congressional races in which the same map was used. The value is calculated using 3 degrees of freedom to create a long-tailed distribution.

The number of competitive seats for a plan is defined as the number of districts whose Democratic and Republican voteshares fall in the 46.5-53.5% range.

We calculate partisan bias by estimating the seat share under a hypothetical tied statewide election, where each party receives exactly 50% of the vote. This measure captures whether the districting plan systematically favors one party in an otherwise evenly split electorate. To do this, we apply a uniform vote swing to all districts so that the statewide average equals 50%. Partisan bias is then defined as the difference between 50% and a party’s seat share. Bias is defined so that a positive value indicates a partisan advantage for Republicans, and a negative value indicates an advantage for Democrats.

In a fair map, both parties would be expected to have similar average win percentages. The packed wins metric represents deviations from this expectation using the difference between the average vote share in each party’s winning districts. If one party is packed into a small number of overwhelmingly favorable districts while being cracked across many unfavorable ones, its average winning margin will be much higher than that of the other party. (Note that statistical significance of this metric can be calculated using the t-test.)

The mean-median difference is the difference between a party’s average vote share and its median vote share across all districts. In a balanced map, the mean and median vote shares should be close to one another. A large difference indicates potential asymmetry in how voters are distributed. If the mean is greater than the median, the party tends to be packed into a few districts with very high vote shares and loses competitiveness elsewhere. (Note that statistical significance of the mean-median difference can be calculated precisely.)

Geography

The geography of a plan is evaluated for its compactness as well as the number of counties split by districts.

We use two metrics to measure compactness,  the Reock and Polsby-Popper scores. Both of these measures range from 0 to 1, where 0 is least compact and 1 is the theoretical maximum. For each plan, all districts are used to calculate the average and minimum Reock and Polsby-Popper scores.

The Reock score is calculated by taking the ratio of the area of the district to the area of the minimum circumscribing circle, or in other words, the smallest circle that entirely encapsulates the district. The Polsby-Popper score is calculated by taking the ratio of the area of the district to the area of the circle whose circumference matches the perimeter of the district, or in other words, the circle that would result by stretching the boundary to form a perfect circle.

Many state constitutions include language to respect existing administrative and political boundaries such as city boundaries and county lines. County splits are measured by counting the number of counties that are divided across districts. We use a more nuanced county-splitting metric developed by Wachspress et al., called split pairs. Split pairs measure the proportion of all possible pairs of voters who live in the same county but are assigned to different districts. A value of 0 means every county is kept whole, while higher values indicate more fragmentation.

Demographics

We use minority voting age population (VAP) to approximate a given minority group’s political influence in a given district. This number is used, in combination with sophisticated analysis of racially polarized voting, in Voting Rights Act legal cases. We include estimates of:

  • Black VAP percentage (BVAP)
  • Hispanic VAP percentage (HVAP)
  • Asian VAP percentage (AVAP)
  • American Indian or Alaska Native VAP percentage (NVAP)
  • Native Hawaiian or Other Pacific Islander VAP percentage (PVAP)

We also sum these to get Minority VAP percentage (MVAP). These metrics allow one to identify majority-minority districts for plans, for purposes of evaluating VRA compliance in particular and racial fairness in general.

Grading

We established a grading scheme to evaluate redistricting plans using the expertise at the Electoral Innovation Lab after inspecting a large number of enacted and hypothetical plans. We then used the same scheme to provide grades and scores for plans across all states in the U.S. This scheme can be improved, and is a subject of ongoing investigation.

We grade each plan according to three categories: partisan fairness, competitiveness, and geographic features. Because partisan fairness is the most robust in detecting gerrymandering harms, it serves as the base grade with the other categories making grade adjustments.

A partisan-fairness grade is assigned as an A, B, C, or F. This serves as a starting point for calculating the overall grade. If competitiveness is an A, the overall grade is improved by one level from the partisan-fairness grade (e.g. B to A). Conversely, if competitiveness is an F, the overall grade is downgraded by one letter (e.g. B to C). The final grade is also downgraded if the geographic features score is an F. Although there is no initial D grade for partisan fairness, a D final grade is achievable by upgrading once from an F or downgrading once from a C.

For each delegation or legislative chamber, a myriad of possible alternative plans are possible. To generate a large set of alternative districting plans which are drawn without partisan intent, we run Markov Chain Monte Carlo simulations on precinct-level statewide files. We have used gerrychain to run these simulations, and currently use the R package redist. These approaches give distributions that are called ensembles.

The ensemble is constrained so that each map must be composed of district population within +/-5% of the average, fewer or the same number of county splits and cut edges as the starting partition, and minority representation as dictated by the backsliding test. The algorithm is run for up to 1,000,000 steps which results in 1,000,000 alternative maps that strictly follow traditional redistricting criteria.

On this ensemble of alternative maps, we calculate the number of Democratic districts and the number of competitive districts using the average of the most recent statewide elections for the U.S. President, U.S. Senate, and Governor. This gives us a distribution of the number of Democratic districts and number of competitive districts for neutrally-drawn maps that could have been drawn in the state that follow traditional criteria. We also calculate geographic metrics to evaluate compactness.

If the number of representatives in a delegation or state chamber has not changed with reapportionment, the last court-accepted plan is used as the starting partition. If the state has a different number of congressional districts than the previous cycle, a random partition is created and then optimized to pass the backsliding test and to have a reasonable number of cut edges (compactness) and county splits relative to the last court-accepted plan.

Partisan Fairness

Partisan fairness is graded by comparing the number of seats produced by a hypothetical plan with a neutrally-defined baseline. The number of seats is evaluated against (a) a range produced by the ensemble of alternative plans as well as (b) an aspirational range of fairness based on the voteshare of a state. Where a plan falls within these ranges determines the partisan fairness grade.

A passing ensemble range is determined by the middle 90% of the distribution of seats won for all simulated plans. That is, a plan passes the ensemble test if it is within the 5th and 95th percentiles of seats.

A passing normative fairness test is based on the statewide vote share. It is calculated using the total number of districts in a state, multiplied by a cumulative normal distribution with a standard deviation of 13.3%. The naturally arising standard deviation of district vote share is an emergent property of how voter preferences vary across precincts, and the specific value of 13.3% corresponds to the the cube law, a historical empirical finding that the ratio of Democratic-to-Republican seat shares, S/(1-S), is proportional to the ratio of the cubes of the vote shares, (V/(1-V))^3. We define the tolerance on either side of this point estimate as the greater of 1 seat or 7% times the total number of seats (5% if there are more than 50 Congressional or 200 legislative seats in total).

When the normative and ensemble ranges overlap, plans are graded according to the following curve:

  • A: passes both normative and ensemble tests
  • B: passes normative test but not ensemble test
  • C: passes ensemble test but not normative test
  • F: fails both normative and ensemble tests

When the normative symmetry range and ensemble range do not overlap, we assign A to the seat share value at the edge of the ensemble range in the direction of the normative range. This seat share represents a naturally occurring result that is also close to being normatively good. We assign B to all seat share values in the normative range and between the normative and ensemble ranges. We assign C to all other seat shares in the ensemble range. Finally, we assign F to all seat shares on the exterior of the ensemble and normative ranges.

Competitiveness

Using the distribution of the number of competitive seats derived from the ensemble, plans are graded according to the following:

  • A: more competitive districts than 95% of the maps in the ensemble
  • C: competitive districts between the 5th and 95th percentile of the distribution
  • F: fewer competitive districts than 95% of the maps in the ensemble

Geographic Features

Using the distribution of average Reock scores derived from the ensemble, plans are graded according to the following:

  • A: higher average Reock score than 95% of the maps in the ensemble
  • B: average Reock score between the 84th and 95th percentile of the distribution
  • C: average Reock score between the 5th and 64th percentile of the distribution
  • F: lower average Reock score than 95% of the maps in the ensemble

Using the distribution of the number of county splits derived from the ensemble, plans are graded according to the following:

  • A: fewer county splits than 95% of the maps in the ensemble
  • C: county splits between the 5th and 95th percentile of the distribution
  • F: more county splits than 95% of the maps in the ensemble

If a map receives an A for both compactness and county splits, it receives an A in geography. If it receives an F for both metrics, it receives an F in geography. If it received one A and one B/C for compactness and county splits, it received a B for geography. All other maps receive a C for geography.

Minority Composition

We use minority voting age population (VAP) to approximate a given minority group’s political influence in a given district. This number is used, in combination with sophisticated analysis of racially polarized voting, in Voting Rights Act legal cases. Though minority composition is not measured in the A-F scale for the final grade, we may override a grade if it fails due to racial gerrymandering. A racial gerrymander is defined using federal standards in place at the time of 2021 redistricting.

A note on the history of the rubric

The report card grading rubric was introduced in 2021. At first, grades were assigned only for states with seven or more districts, and metrics-only for smaller delegations. Starting in 2023, the grading system was modified to allow any plan with two or more districts to be graded.

In the first round of grading (2021-2022), ensemble methods used the Python library GerryChain to simulate alternative redistricting plans. In 2023 we started to use redist, which allows for easier integration of data into the simulation and scoring pipeline. The results are very similar, and in our spotchecks, ensemble seat ranges are not substantially affected.

We originally measured seat share on an all-or-none basis for each seat. To compensate for the integer-resolution of this approach we gave plans some leeway: if a plan fell inside the ensemble range but outside the normative range, we extended the normative range by 7% on each end and awarded a “B” if the plan fit within this expanded range. After improving our calculation, which now uses real numbers of seats, we now no longer apply the leeway extension.

For the principal wave of redistricting in 2021 and 2022, competitiveness, compactness, and county splits were all graded using a three-step curve, A-C-F. Starting in 2023, we added a B range for compactness by adding a cut at the 84th percentile.

Page last updated:Jul 19th 2026