Political parties in India are federated organisations with a persistent knowledge problem. The central office wants to know how each district unit is doing — whether the local committee is functional, whether it holds meetings, whether it has surfaced credible candidates, whether the campaign infrastructure exists in time for the next cycle. That knowledge sits with the observer the party sends to look, and it sits in the reports that observer files.
The reports themselves are a strange artefact. They are usually PDFs uploaded through a portal. They contain candidate names, meeting attendance, sometimes photographs, sometimes just a signature and a date. They accumulate in a database that nobody reads.
The data project we ran turned that accumulation into a monthly grade — one score per observer, in one of four bands, across more than seven hundred districts in twenty-three states. This is a note on what the exercise measured, what it refused to measure, and what happened when the numbers were shown to the people they described.
The seven parameters
The framework scores each observer on seven parameters, each with a defined weight, each drawn from something already sitting in the database. It looks like this.
Report quality. Whether the report is completed, whether the fields have real content or the same paragraph repeated across three districts, whether the photographs are legible, whether the dates are internally consistent. The measure is boring by design. Boring reports get filed. Non-boring reports get flagged.
Daily activity. Whether the observer’s logged activity has any signal at all across the month. The failure mode this catches is the observer who files everything on the last day, at 11:47pm, with the same timestamp on twelve district reports.
Documentation. Whether the supporting attachments — meeting minutes, photographs, candidate resumes — are present, and whether they cross-reference correctly. If the report says a meeting happened on the tenth and the attached photograph has a metadata timestamp from the seventeenth, the score drops. Not because either is disqualifying, but because the mismatch is a signal worth reading.
Proposed candidates. Whether the observer has surfaced names, whether those names carry the required biographical detail, whether the demographic mix of the proposed pool matches the district’s population. This is the parameter that does the real political work; more on that in a minute.
Timeliness. Whether the report arrived before the reporting window closed. Late reports drop a full band. Missing reports drop two.
Cross-verification. Whether the observer’s claims about attendance, resources, and candidate viability hold up against the two nearest data sources — the district committee’s own reports, and the state secretariat’s activity logs. Divergence flags for review.
Signal-to-noise. A rolling metric that watches for anomalies in each observer’s own history. A sudden change in reporting style, a lull followed by a burst, a run of grade-A reports from a district known to be dormant — all get flagged. The metric is calibrated to the observer, not the population, so it catches within-person changes that a cross-sectional analysis would miss.
The bias diagnostic
The seven parameters run on the observer. A separate diagnostic runs on the observer’s choices — specifically, on the demographic composition of the candidates they rank first across the districts they cover.
Two patterns fell out that the party would not have found by reading reports one by one.
One. In districts where the observer’s own caste was one of two dominant local castes, ninety-plus percent of first-preference candidates came from those two castes. That is not necessarily disqualifying — dominant-caste candidates may genuinely have the strongest local base — but the pattern is uniform across observers of that background, and it disappears when the analysis is repeated in observers whose caste is not dominant locally. The party can decide what to do with this. What it cannot do, after the analysis, is pretend it did not know.
Two. In several majority-Adivasi districts, first-preference candidates included zero Scheduled Tribe names. This surfaced only because the diagnostic compared the caste mix of the proposed pool against the census demographics of the district. Without the automatic comparison, the pattern lives in the observer’s inbox indefinitely.
Three, adjacent to both. Proposer phone numbers occasionally matched candidate phone numbers. The proposer is required to be an independent local leader; the candidate is the person being proposed. When the two phone numbers are the same person, the process has collapsed on itself. The diagnostic flagged eleven such cases across the first quarter. Each one deserves a human decision. The point is that the human decision now has data to make.
Why the score has to be visible
The score, on its own, would have been an academic exercise. What made the data project useful was that the grade a district’s observer received was returned to that observer, along with a per-parameter breakdown, at the end of each monthly cycle. Not as an accusation. As a mirror.
The behavioural response was uneven and quick. Reports got longer. Attachments got attached. Timestamps started to spread across the month instead of clustering on the last day. Some observers escalated to the state office to argue with their grade, which is exactly what should happen — an argument about the grade is an argument about what the grade measures, and that is a productive place for the conversation to be.
The metric it moved slowest on was the caste-representation diagnostic. That one requires a political intervention, not a documentary one, and no scoring framework can substitute for the decision to change the pool.
What the framework is not
The framework is not an audit tool. Audits are done under different constraints, by different people, and with different powers to sanction. What this is, is an operating diagnostic — a monthly reading that surfaces what the reports collectively say, so the party can act on the pattern before it becomes a headline.
The distinction matters because when the exercise gets described as an audit, the political reflex is defensive. When it gets described as a mirror the observer chose to hold up, the reflex is different. The name of the tool, in this kind of work, is part of the tool.
If you run a similar exercise in a differently-organised institution and it does something we did not model here, write in. The framework is deliberately generic across the seven parameters, and any of them can be replaced with something that fits your data.