Skip to content

v0.10.0

A release about saying what the data is. Two new mechanisms record things the library had no way to express -- what a sample represents, and whether a column that is present actually holds anything -- and one behaviour change makes sampling weights comparable across waves. Alongside those, a long run of country correctness fixes, several of which recovered data that had been silently absent for years.

⚠️ Your cache rebuilds once on first use

LSMS_CACHE_SCHEMA stays at 5, but two changes edit get_dataframe / read_file, and df_data_grabber is @build_transform()-tagged and references them -- so every table's lsms_cache_hash moves: the null-read guard's read-path site (#711) and the .DAT/.DCT layout detection (#708). 37 of 37 probed hashes move, and a bare no-op stub in get_dataframe moves the same 37 -- the invalidation is a property of touching that function, not of the guard. Both ship together so the rebuild happens once. No action needed.

Behaviour change: weights are normalised to within-wave mean 1

sample().weight and sample().panel_weight now come back with mean 1.0 in every wave, normalised at API time in Country._finalize_result (#718). The parquets keep the raw values -- a read-path transform like kinship expansion, so pd.read_parquet(cache_path) still shows expansion weights.

The corpus shipped two incompatible scales with nothing marking which was which: of 93 weighted (country, wave) cells across 31 countries, 13 were already normalised and 80 were expansion weights summing to a national population (wave means 25.7 to 8,659). CotedIvoire held both across its own waves, so a user summing across them silently got numbers three orders of magnitude apart; GhanaLSS had avoided the hazard by leaving four waves' weights NULL. The rule needs no threshold and no configuration -- divide each column by its own wave's non-null mean. After it, 93 of 93 weight cells and 83 of 83 panel_weight cells sit at mean 1 (max deviation 4.44e-16).

Preserved: weighted ratios, shares and means, exactly -- scaling every weight in a wave by one positive constant cancels in Sw*x / Sw. Measured against the raw parquet over 21 waves in four countries, worst absolute difference 2.2e-16.

Not preserved: sum(weight) is now a wave's household count, not a population estimate, and cross-wave pooling reflects sample size rather than population. On a pooled weighted Rural share that moves CotedIvoire by +0.0349, Malawi +0.0083, Uganda +0.0068, GhanaLSS -0.0078. A miscoded weight variable can no longer be caught by a cross-wave population jump at API level; that check moved to the raw parquet layer, where scale still exists.

All seven GhanaLSS waves now carry weights (#713, #714). The null-read guard's first cold run reported weight 100% NULL in GLSS1-GLSS4. GLSS4 is wired to real varying weights from POV_GH.DTA (298 distinct values survive normalisation -- GSS states GLSS4 is not self-weighting); GLSS3 relays GSS's shipped constant; GLSS1/GLSS2 are asserted at 1.0 from the report's self-weighting statement, and the docstrings say in those words that relaying a producer value and asserting one are different acts. weight non-null goes 39,468 -> 56,330 of 56,330.

New: the population record

Two "nationally representative" surveys can represent different populations. countries/{C}/_/population.yml records, per wave, what the survey's own documentation says its sample represents (#720, closes #603; refs #601) -- 116 wave records across 38 countries in ten universe tags (national-all-households 39, national-claimed 29, region-excluded 19, subnational-area 11, not-stated 7, panel-inherited 5, and five more). Read it via Country(name).population or off df.attrs['population']. Guide: docs/guide/population.md.

It warns; it never fences. Feature() returns exactly what it returned before -- a default that silently drops data is the same disease as one that silently pools it. PopulationHeterogeneityWarning fires once per call, names the classes in the pool, and says nothing was dropped. A real Feature('sample')() (534,037 rows, 34 countries) emits exactly one warning, naming 8 materially different universes plus 6 waves whose universe no document states.

universe_tag never travels alone: PopulationRecord.from_config refuses a record missing source_type or confidence, because the tag is an editorial reading and confidence: low means "this is sample-design text and should not be laundered into a universe". unrecorded (nobody looked) and not-stated (the sweep looked, and found nothing) stay distinct. The files are generated from slurm_logs/POPULATION_STATEMENTS_2026-07-21.org by scripts/promote_population_records.py. population.yml is a sibling of data_scheme.yml, never a block inside it: all 524 (country, table) cache hashes are byte-identical before and after.

df.attrs survival, re-measured under pandas 3.0.2. The per-method rule ("merge drops attrs") is wrong in both directions. One rule covers it: attrs survive an operation only when every input agrees; any disagreement, including one side having none, yields {}. TestAttrsSurvival pins all seven cells, including the preserving ones, so a future pandas that dropped unconditionally goes red rather than making the explicit re-attach in _join_v_from_sample look redundant.

New: the null-read guard

Every guard the library had checked shape -- a declared column is present, an index is unique. None looked inside the column, so a parse returning a correctly-shaped frame of NaN was served silently. lsms_library/null_read_audit.py adds two sites (#711): Site R in get_dataframe, catching the mis-parse class including columns nothing has wired yet (a wave is usually mis-parsed before anyone wires it); and Site B in Country._finalize_result, catching a required declared column that is present and empty. Niger 2014-15's Latitude (0 of 270) is invisible to Site R -- that wave declares none, so no read produces it.

The trigger is a fraction of the frame, not "any all-null column", and the measurement is why. Over 2,264 declared reads / 65,099 columns swept cold, the per-column test fires 887 times and only 3 of those are columns a table asked for: a firehose, and a warning nobody reads is how #323 survived its first fix. The frame-fraction test fires 0 times, with a corpus ceiling of 28.3% against 40-82% for every known-bad parse, so _NULL_FRACTION_TRIGGER = 1/3 sits in a measured gap. Site B over 504 cold builds: 0 required columns null in every wave, 86 null in some wave's t slice across 42 tables and 17 countries -- a work queue, and the first thing it found was GhanaLSS's four weightless waves.

Warn by default; LSMS_READ_STRICT=1 makes it fatal (NullReadError). Deliberately its own lever, not LSMS_GRAIN_STRICT: a destroyed row and an empty column are different concerns and must ratchet separately. null_read_reports() is the public accessor, twin of grain_reports(). No allowlist. Overhead ~3% on the corpus's largest read, 0.1% on the largest built table.

Readers and access paths

.DAT layout is decided by its .DCT, not assumed (#708). get_dataframe reached a bare read_csv, right about the delimiter and wrong about the rest: . is Stata's ASCII missing marker and nothing honoured it, so one missing value demoted a numeric column to str and chunked inference then produced mixed int/str object columns. GhanaLSS 1987-88 Y12A.DAT:FOODCD -- the food_acquired item code j -- has 61 real codes and reported 123 distinct values, and a groupby on such a column splits every key in two. 449 of 531 dictionary-paired files gain correctly-typed columns, and four Nicaragua files that were hard-unreadable now read. Detection is three ordered tests with no thresholds and no tie-breaks, otherwise DatLayoutError; the sweep of all 535 .DAT files finds 531 csv-header, 4 with no dictionary, 0 unknown. Three GhanaLSS consumers the string dtype had been hiding are repaired in the same PR.

pandas >= 3.0 is now a floor (#710, refs #709). 1.7498005798264095e+100 is exactly 2^333 -- Stata's system-missing . for a double in .dta format 104/105, the third and oldest sentinel. pandas nulls it as of 3.0; 2.2 returns a finite float. Ten country-waves ship 104/105 files (~28.5M reader-confirmed cells): on 2.2, GhanaLSS 1991-92 food_acquired.Price has mean 1.71e99 against a correct 342.6, and Panama 1997 Expenditure mean 1.31e105 against 7.65e7. A data-correctness floor, not an API one -- the code runs fine on 2.2 -- and a user-facing exclusion of installs pinned below 3.0.

get_data_file accepts repo-relative paths (#717, closes #716), like get_dataframe already did. It had been building .../countries/lsms_library/countries/... and returning None, which propagated as the string 'None' and raised FileNotFoundError: 'None' pointing at the wrong tool. That misdirected the #713 investigation twice.

Country data correctness

Nigeria food_prices was dropping ~99% of its rows (#669, closes #591) -- 675 to 1,106 rows against a food_acquired of 120k-160k, in 6 of 8 waves, with no warning and no NaN. Not a sentinel: q5a is consumed out of purchases, not own production (that is q6a), and q10 pays for q9a in unit q9b. The coverage matrix graded all of those cells sane -- a direct illustration of why sane is not blessed. The durable half is country-agnostic: UnpriceableRowsWarning fires when more than 5% of rows that should have produced a price do not, naming both candidate causes.

A validator that never ran (#671, closes #537). validate_acquisition_source had zero call sites since its introducing commit while its docstring asserted it was called from _finalize_result. Now wired in; NaN counts as a violation (it previously called .dropna() first, blinding it to the dominant real failure); and it reads unique values off MultiIndex.levels -- 998 ms to 2.8 ms on a 5.26M-row frame.

Ghana: 56,817 people had no kinship data (#673). get_categorical_mapping called without a value-column keyword returns an empty dict -- no exception, every lookup misses, the column ships 100% null and the build reports success. An AST sweep of all 77 call sites found exactly 7 broken, all GhanaLSS, in three waves; fixed with a new additive local_tools.code_label_map. Reviving the decodes then armed two latent errors that were harmless while the columns were empty: GLSS4's region table had codes 9 and 10 swapped, and GLSS1 fell through to a country-level table with Upper East/West reversed -- wrong for 1,511 people and 11 of 176 clusters (#676). A fallback that has never been exercised has never been checked. household_roster is now complete in all seven waves.

#323 country configs: Uganda, Albania, Togo, Guyana, Serbia, China, Pakistan, GhanaLSS and Niger. All config-only; none adds an aggregation: key, because a duplicate on a declared index means the identifier is broken.

  • Pakistan (#612) built cluster_features from the household cover sheet -- 4,957 rows into a 301-cluster table, 4,656 silently discarded, Region wired to religion and Language to the language of the interview; 1,101 households had been given a cluster value contradicting their own record.
  • Uganda (#634) people_last7days went long-form in 2018-19 with the member/visitor key commented out, so first() took whichever row came first: 1,636 of 3,242 households were served visitor counts. Mean people per household 2.81 -> 4.51, against 5.0-5.4 in every earlier wave.
  • Serbia (#666): popkrug is local to a municipality, so 510 districts shared 328 values and 182 clusters vanished. Guyana (#668, closes #503): 257 (ED, HH) buckets fused 562 distinct households holding 2,419 of 7,827 people; COVERN.NEWID proves the key is (ED, SN, HH) on 1807/1807 rows.
  • Albania (#620): 2004's m0_distr is a household's current district while the PSU code is its original one, so movers set whole clusters' districts -- cluster 223 was decided by a lone mover.
  • China (#667) extracted a village-level table at person grain (3,002 rows for 30 villages); Togo (#632) handed a 540-grappe table 6,171 households.
  • GhanaLSS (#622): birthplace-as-cluster-region is unwired (its code list includes Nigeria and Togo), and 2016-17 food_security is rekeyed on hid, recovering 110 households from one phantom (t, NaN) tuple.

Malawi 2016-17 (#672): four YAML-path tables declared a bare i: case_id and emitted ids in a different namespace from the rest of the wave -- 48,589 rows with 0% overlap against household_roster, now 100%.

Key-soundness review of the remaining #637 groupby().first() sites (#646,

651): Uganda plot_inputs needed a season level, and Tanzania 2008-15

shocks was keyed on the panel line (UPHI), not the household. The other sites reviewed are sound.

Peru's config was unreachable (#691, closes #684): its Data Scheme: block sat in a file named data_info.yml, which Country.resources never reads, so Country('Peru').data_scheme was [] and Peru contributed zero cells to the coverage matrix -- a country whose config does not load cannot even be graded broken. Renamed and wired: 364 clusters, 19,285 people, 3,621 households, zero GrainCollapseWarning. TestCountryConfigIsReachable pins it.

Rural and urban definitions

A corpus-wide audit of every cell producing Rural (#682) measured 150 of 153 producing cells at source and cold-delivered all 47 countries: 12 cells collapse a multi-way source onto two values, 11 are declared but >=99% null, and 0 either silently null a raw value or deliver a non-canonical one. Then:

  • GLSS1/GLSS2 cluster Region is measured, not inferred (#686, closes #681). It had been the modal birth region of a cluster's under-12s. BID Appendix I -- the survey's own cluster list -- is transcribed to _/appendix_i_clusters.org (263 rows; 176/176 and 170/170 clusters matched), and five clusters move to the adjacent region the inference had put them next to.
  • GLSS1/GLSS2 Rural nulls 6,341 -> 0 (#690, closes #685). Appendix I is three-way (U/R/SU); SU is delivered as Rural and no canonical Semi-urban is added, so lsms_library/data_info.yml is unchanged. The fold is a judgement call the survey does not settle, and it is not recoverable from the delivered column -- the raw three-way is kept in the org table.
  • GLSS5 Rural nulls 8,687 -> 0 (#687, closes #659) by wiring loc2 from parta/sec7.dta. Two decoys are recorded in the config so nobody re-picks them, one carrying loc2 present and 100% null.
  • Ethiopia's Rural is not comparable across waves (#688, closes #683). Urban goes 12.7% -> 54.0% of the sample from ESS1 to ESS4 while the rural count is flat (3,466 -> 3,115): the sample is being extended into urban strata, not shifted, and ESS1 was designed for rural areas and small towns. Documentation only.
  • Settlement thresholds are scoped (#704): what defeats cross-country settlement harmonisation is variation in the entity classified (locality vs enumeration area vs commune), not the population threshold.

New coverage

GhanaSPS gains four tables: sample (#673 -- data_scheme.yml had asserted "cluster identity unavailable", which was false; 334 EAs in all three waves), household_roster (#677 -- 54,251 people, and household_characteristics with it), housing (#726) and individual_education (#728). Each records the trap that would have made it silently wrong: the module letter shifts between waves while the variable names do not, so wave 3's 01fi_* is employment where wave 2's is education.

livestock is defined canonically, and its value column split by additivity (#736). There was no canonical block, so three different quantities were all spelled Value and all graded sane: Feature('livestock') had been stacking 71,121 per-head prices (Uganda, Nigeria, Malawi) on 23,514 sales totals (Ethiopia, Mali) in one column. Split by what arithmetic is legal -- ValuePerAnimal (not additive), HerdValue (a stock), SalesValue (a flow) -- the discriminating test being whether the value is reported when the household sold nothing, settled against questionnaire wording. Values are unchanged: rows, nulls, sums and describe() identical for all 14 countries. tests/test_currency_livestock_registry.py pins currency.py's second copy of "which columns hold money" against the canonical YAML by set equality; a stale entry there has no symptom at all, and silently leaves a renamed column unconverted.

The coverage matrix grades against a pinned feature vocabulary (#725, closes #724). build_matrix took its feature axis from each country's own config, so a feature a country never declared produced no row -- GhanaSPS reported 21/21 = 100% sane while ~300 of its 328 source files sat unwired. That pre-closed 417 (country, feature) pairs as "not applicable" with no evidence and no record of who decided. The 22-feature vocabulary is measured, not tuned: ranking the 33 source features by declaring-country count leaves a real gap at 22 (nothing has 3, 4 or 5), so >=3 and >=6 select the identical set. The shipped snapshot goes 1,849 -> 2,279 rows with 415 undeclared cells; a not-asked or asked-not-distributed verdict closes one, so it is a queue, not 415 permanently-red cells. It also fixes a pre-existing snapshot bug: --no-readiness had been silently downgrading 1,217 sane cells to declared in the committed snapshot.

labels= accepts a dict (#697, refs #682/#685) -- c.sample(labels={'Rural': 'Settlement'}) selects a label column for any mapped column or index level. The scalar form still targets j only and is byte-for-byte unchanged. Selection happens at the mapping site, keyed on Code: a settlement ladder is coarse-to-fine, so keying on Preferred Label gives two arbitrary entries for seven rungs. No country config is wired; this is the mechanism. Cache hashes are unchanged, and a structural test walks the real @build_transform closure so a refactor that drags the label path into the build fingerprint fails a test rather than invalidating the corpus. Guide: docs/guide/label-selection.md.

Tests, CI and documentation

  • --collect-only no longer purges the cache (#722, closes #719) -- collection is a dry run. pytest collection is scoped to tests/ (#733), after a permutation test shipped alongside a findings doc did its data loading at module level and aborted the whole CI session, naming neither a test nor the change under review. A guard checked with ast, not text matching, catches a dependency shadowing the tests package (#687 carried this alongside the GLSS5 fix above; refs #680).
  • docs/guide/population.md and docs/guide/label-selection.md are new; coverage.md gains the undeclared tier and caching.md records weights as the clearest parquet-vs-API divergence.
  • Seventeen task ledgers under .coder/ledger/, and 11 GhanaLSS absent_verdicts rows closing 6 cells (#678) -- Ghana previously had zero.
  • GhanaSPS is not a post-planting/post-harvest country (#735, closes #730): no wave ships two rounds. Its constructible agriculture features are sized as todo (#729), and plotid is documented as not a panel id after wave 2 -- tested, negative.
  • A branch audit records why 26 #323 country fixes had never merged (#670).

Known open

  • #323 and #637 remain open. The country fixes here are a large block of them, not the whole class.
  • #709: Kyrgyz 1998 SECT_06.DTA:fprimary = -1.2255e294 is a distinct value pandas does not filter, and it is delivered. The pandas 3.0 floor closes the 2^333 exposure only.
  • LSMS_READ_STRICT=1 is not on in CI. optional: is country-grain while a null-read finding is wave-grain, so there is no way to record that Niger 2014-15's missing coordinates are correct without disarming the column for every wave. That gates the ratchet.
  • GLSS2's asserted weight of 1.0 assumes attrition from the GLSS1 rotating panel was ignorable, which nobody has shown. PANELC.DAT carries the linkage if a genuine panel weight is ever built.
  • #592: 18 waves across 6 countries still have zero GPS in cluster_features.
  • EthiopiaRHS livestock and income are misdated (#738). The seven {village}lvs5.tab and four {village}inc5.tab files sit in the 1999/ wave directory and are served at t='1999', but they are the 1989 IFPRI survey -- convergent evidence from SPSS _date stamps, the packed 1989 HHID key, 95-98% household overlap with the 1989-wired asset{v}.tab siblings against zero overlap with the R5 roster, and a village set that is exactly 1989's seven PAs. 541 households are affected; 2004 and 2009 rows are not. Documented, deliberately not rewired -- re-dating a wave touches sample / cluster_features / panel_ids and the coverage matrix. This is the one country where the canonical livestock schema above sits on top of data carrying the wrong t.

Upgrading

  1. pandas >= 3.0 is required. Installs pinned below it are excluded; see #710 for why.
  2. The first build after upgrading is a cold one. Nothing to do.
  3. If you use sample().weight, check whether you relied on its scale. Weighted ratios, shares and means are unchanged to machine precision; sum(weight) is now a household count. If you need expansion weights, read the L2-country parquet, which keeps the raw values.