v0.10.0¶
A release about saying what the data is. Two new mechanisms record things the library had no way to express -- what a sample represents, and whether a column that is present actually holds anything -- and one behaviour change makes sampling weights comparable across waves. Alongside those, a long run of country correctness fixes, several of which recovered data that had been silently absent for years.
⚠️ Your cache rebuilds once on first use¶
LSMS_CACHE_SCHEMA stays at 5, but two changes edit get_dataframe /
read_file, and df_data_grabber is @build_transform()-tagged and
references them -- so every table's lsms_cache_hash moves: the null-read
guard's read-path site (#711) and the .DAT/.DCT layout detection (#708).
37 of 37 probed hashes move, and a bare no-op stub in get_dataframe moves
the same 37 -- the invalidation is a property of touching that function, not of
the guard. Both ship together so the rebuild happens once. No action needed.
Behaviour change: weights are normalised to within-wave mean 1¶
sample().weight and sample().panel_weight now come back with mean 1.0 in
every wave, normalised at API time in Country._finalize_result (#718). The
parquets keep the raw values -- a read-path transform like kinship expansion,
so pd.read_parquet(cache_path) still shows expansion weights.
The corpus shipped two incompatible scales with nothing marking which was
which: of 93 weighted (country, wave) cells across 31 countries, 13 were
already normalised and 80 were expansion weights summing to a national
population (wave means 25.7 to 8,659). CotedIvoire held both across its own
waves, so a user summing across them silently got numbers three orders of
magnitude apart; GhanaLSS had avoided the hazard by leaving four waves' weights
NULL. The rule needs no threshold and no configuration -- divide each column by
its own wave's non-null mean. After it, 93 of 93 weight cells and 83 of 83
panel_weight cells sit at mean 1 (max deviation 4.44e-16).
Preserved: weighted ratios, shares and means, exactly -- scaling every
weight in a wave by one positive constant cancels in Sw*x / Sw. Measured
against the raw parquet over 21 waves in four countries, worst absolute
difference 2.2e-16.
Not preserved: sum(weight) is now a wave's household count, not a
population estimate, and cross-wave pooling reflects sample size rather than
population. On a pooled weighted Rural share that moves CotedIvoire by
+0.0349, Malawi +0.0083, Uganda +0.0068, GhanaLSS -0.0078. A miscoded
weight variable can no longer be caught by a cross-wave population jump at API
level; that check moved to the raw parquet layer, where scale still exists.
All seven GhanaLSS waves now carry weights (#713, #714). The null-read
guard's first cold run reported weight 100% NULL in GLSS1-GLSS4. GLSS4 is
wired to real varying weights from POV_GH.DTA (298 distinct values survive
normalisation -- GSS states GLSS4 is not self-weighting); GLSS3 relays GSS's
shipped constant; GLSS1/GLSS2 are asserted at 1.0 from the report's
self-weighting statement, and the docstrings say in those words that relaying a
producer value and asserting one are different acts. weight non-null goes
39,468 -> 56,330 of 56,330.
New: the population record¶
Two "nationally representative" surveys can represent different populations.
countries/{C}/_/population.yml records, per wave, what the survey's own
documentation says its sample represents (#720, closes #603; refs #601) --
116 wave records across 38 countries in ten universe tags
(national-all-households 39, national-claimed 29, region-excluded 19,
subnational-area 11, not-stated 7, panel-inherited 5, and five more).
Read it via Country(name).population or off df.attrs['population']. Guide:
docs/guide/population.md.
It warns; it never fences. Feature() returns exactly what it returned
before -- a default that silently drops data is the same disease as one that
silently pools it. PopulationHeterogeneityWarning fires once per call, names
the classes in the pool, and says nothing was dropped. A real
Feature('sample')() (534,037 rows, 34 countries) emits exactly one warning,
naming 8 materially different universes plus 6 waves whose universe no document
states.
universe_tag never travels alone: PopulationRecord.from_config refuses a
record missing source_type or confidence, because the tag is an editorial
reading and confidence: low means "this is sample-design text and should not
be laundered into a universe". unrecorded (nobody looked) and not-stated
(the sweep looked, and found nothing) stay distinct. The files are generated
from slurm_logs/POPULATION_STATEMENTS_2026-07-21.org by
scripts/promote_population_records.py. population.yml is a sibling of
data_scheme.yml, never a block inside it: all 524 (country, table) cache
hashes are byte-identical before and after.
df.attrs survival, re-measured under pandas 3.0.2. The per-method rule
("merge drops attrs") is wrong in both directions. One rule covers it:
attrs survive an operation only when every input agrees; any disagreement,
including one side having none, yields {}. TestAttrsSurvival pins all
seven cells, including the preserving ones, so a future pandas that dropped
unconditionally goes red rather than making the explicit re-attach in
_join_v_from_sample look redundant.
New: the null-read guard¶
Every guard the library had checked shape -- a declared column is present,
an index is unique. None looked inside the column, so a parse returning a
correctly-shaped frame of NaN was served silently.
lsms_library/null_read_audit.py adds two sites (#711): Site R in
get_dataframe, catching the mis-parse class including columns nothing has
wired yet (a wave is usually mis-parsed before anyone wires it); and Site B
in Country._finalize_result, catching a required declared column that is
present and empty. Niger 2014-15's Latitude (0 of 270) is invisible to Site R
-- that wave declares none, so no read produces it.
The trigger is a fraction of the frame, not "any all-null column", and the
measurement is why. Over 2,264 declared reads / 65,099 columns swept cold,
the per-column test fires 887 times and only 3 of those are columns a
table asked for: a firehose, and a warning nobody reads is how #323 survived
its first fix. The frame-fraction test fires 0 times, with a corpus ceiling
of 28.3% against 40-82% for every known-bad parse, so
_NULL_FRACTION_TRIGGER = 1/3 sits in a measured gap. Site B over 504 cold
builds: 0 required columns null in every wave, 86 null in some wave's
t slice across 42 tables and 17 countries -- a work queue, and the first
thing it found was GhanaLSS's four weightless waves.
Warn by default; LSMS_READ_STRICT=1 makes it fatal (NullReadError).
Deliberately its own lever, not LSMS_GRAIN_STRICT: a destroyed row and an
empty column are different concerns and must ratchet separately.
null_read_reports() is the public accessor, twin of grain_reports(). No
allowlist. Overhead ~3% on the corpus's largest read, 0.1% on the largest built
table.
Readers and access paths¶
.DAT layout is decided by its .DCT, not assumed (#708). get_dataframe
reached a bare read_csv, right about the delimiter and wrong about the rest:
. is Stata's ASCII missing marker and nothing honoured it, so one missing
value demoted a numeric column to str and chunked inference then produced
mixed int/str object columns. GhanaLSS 1987-88 Y12A.DAT:FOODCD -- the
food_acquired item code j -- has 61 real codes and reported 123 distinct
values, and a groupby on such a column splits every key in two. 449 of 531
dictionary-paired files gain correctly-typed columns, and four Nicaragua files
that were hard-unreadable now read. Detection is three ordered tests with no
thresholds and no tie-breaks, otherwise DatLayoutError; the sweep of all 535
.DAT files finds 531 csv-header, 4 with no dictionary, 0 unknown. Three
GhanaLSS consumers the string dtype had been hiding are repaired in the same
PR.
pandas >= 3.0 is now a floor (#710, refs #709). 1.7498005798264095e+100
is exactly 2^333 -- Stata's system-missing . for a double in .dta format
104/105, the third and oldest sentinel. pandas nulls it as of 3.0; 2.2 returns
a finite float. Ten country-waves ship 104/105 files (~28.5M reader-confirmed
cells): on 2.2, GhanaLSS 1991-92 food_acquired.Price has mean 1.71e99
against a correct 342.6, and Panama 1997 Expenditure mean 1.31e105
against 7.65e7. A data-correctness floor, not an API one -- the code runs fine
on 2.2 -- and a user-facing exclusion of installs pinned below 3.0.
get_data_file accepts repo-relative paths (#717, closes #716), like
get_dataframe already did. It had been building
.../countries/lsms_library/countries/... and returning None, which
propagated as the string 'None' and raised FileNotFoundError: 'None'
pointing at the wrong tool. That misdirected the #713 investigation twice.
Country data correctness¶
Nigeria food_prices was dropping ~99% of its rows (#669, closes #591) --
675 to 1,106 rows against a food_acquired of 120k-160k, in 6 of 8 waves, with
no warning and no NaN. Not a sentinel: q5a is consumed out of purchases,
not own production (that is q6a), and q10 pays for q9a in unit q9b.
The coverage matrix graded all of those cells sane -- a direct
illustration of why sane is not blessed. The durable half is
country-agnostic: UnpriceableRowsWarning fires when more than 5% of rows that
should have produced a price do not, naming both candidate causes.
A validator that never ran (#671, closes #537).
validate_acquisition_source had zero call sites since its introducing commit
while its docstring asserted it was called from _finalize_result. Now wired
in; NaN counts as a violation (it previously called .dropna() first, blinding
it to the dominant real failure); and it reads unique values off
MultiIndex.levels -- 998 ms to 2.8 ms on a 5.26M-row frame.
Ghana: 56,817 people had no kinship data (#673).
get_categorical_mapping called without a value-column keyword returns an
empty dict -- no exception, every lookup misses, the column ships 100% null and
the build reports success. An AST sweep of all 77 call sites found exactly 7
broken, all GhanaLSS, in three waves; fixed with a new additive
local_tools.code_label_map. Reviving the decodes then armed two latent
errors that were harmless while the columns were empty: GLSS4's region table
had codes 9 and 10 swapped, and GLSS1 fell through to a country-level table
with Upper East/West reversed -- wrong for 1,511 people and 11 of 176 clusters
(#676). A fallback that has never been exercised has never been checked.
household_roster is now complete in all seven waves.
#323 country configs: Uganda, Albania, Togo, Guyana, Serbia, China,
Pakistan, GhanaLSS and Niger. All config-only; none adds an aggregation: key,
because a duplicate on a declared index means the identifier is broken.
- Pakistan (#612) built
cluster_featuresfrom the household cover sheet -- 4,957 rows into a 301-cluster table, 4,656 silently discarded,Regionwired toreligionandLanguageto the language of the interview; 1,101 households had been given a cluster value contradicting their own record. - Uganda (#634)
people_last7dayswent long-form in 2018-19 with the member/visitor key commented out, sofirst()took whichever row came first: 1,636 of 3,242 households were served visitor counts. Mean people per household 2.81 -> 4.51, against 5.0-5.4 in every earlier wave. - Serbia (#666):
popkrugis local to a municipality, so 510 districts shared 328 values and 182 clusters vanished. Guyana (#668, closes #503): 257(ED, HH)buckets fused 562 distinct households holding 2,419 of 7,827 people;COVERN.NEWIDproves the key is(ED, SN, HH)on 1807/1807 rows. - Albania (#620): 2004's
m0_distris a household's current district while the PSU code is its original one, so movers set whole clusters' districts -- cluster 223 was decided by a lone mover. - China (#667) extracted a village-level table at person grain (3,002 rows for 30 villages); Togo (#632) handed a 540-grappe table 6,171 households.
- GhanaLSS (#622): birthplace-as-cluster-region is unwired (its code list
includes Nigeria and Togo), and 2016-17
food_securityis rekeyed onhid, recovering 110 households from one phantom(t, NaN)tuple.
Malawi 2016-17 (#672): four YAML-path tables declared a bare i: case_id
and emitted ids in a different namespace from the rest of the wave -- 48,589
rows with 0% overlap against household_roster, now 100%.
Key-soundness review of the remaining #637 groupby().first() sites (#646,
651): Uganda plot_inputs needed a season level, and Tanzania 2008-15¶
shocks was keyed on the panel line (UPHI), not the household. The other
sites reviewed are sound.
Peru's config was unreachable (#691, closes #684): its Data Scheme: block
sat in a file named data_info.yml, which Country.resources never reads, so
Country('Peru').data_scheme was [] and Peru contributed zero cells to the
coverage matrix -- a country whose config does not load cannot even be graded
broken. Renamed and wired: 364 clusters, 19,285 people, 3,621 households,
zero GrainCollapseWarning. TestCountryConfigIsReachable pins it.
Rural and urban definitions¶
A corpus-wide audit of every cell producing Rural (#682) measured 150 of 153
producing cells at source and cold-delivered all 47 countries: 12 cells
collapse a multi-way source onto two values, 11 are declared but >=99% null,
and 0 either silently null a raw value or deliver a non-canonical one. Then:
- GLSS1/GLSS2 cluster
Regionis measured, not inferred (#686, closes #681). It had been the modal birth region of a cluster's under-12s. BID Appendix I -- the survey's own cluster list -- is transcribed to_/appendix_i_clusters.org(263 rows; 176/176 and 170/170 clusters matched), and five clusters move to the adjacent region the inference had put them next to. - GLSS1/GLSS2
Ruralnulls 6,341 -> 0 (#690, closes #685). Appendix I is three-way (U/R/SU);SUis delivered asRuraland no canonicalSemi-urbanis added, solsms_library/data_info.ymlis unchanged. The fold is a judgement call the survey does not settle, and it is not recoverable from the delivered column -- the raw three-way is kept in the org table. - GLSS5
Ruralnulls 8,687 -> 0 (#687, closes #659) by wiringloc2fromparta/sec7.dta. Two decoys are recorded in the config so nobody re-picks them, one carryingloc2present and 100% null. - Ethiopia's
Ruralis not comparable across waves (#688, closes #683). Urban goes 12.7% -> 54.0% of the sample from ESS1 to ESS4 while the rural count is flat (3,466 -> 3,115): the sample is being extended into urban strata, not shifted, and ESS1 was designed for rural areas and small towns. Documentation only. - Settlement thresholds are scoped (#704): what defeats cross-country settlement harmonisation is variation in the entity classified (locality vs enumeration area vs commune), not the population threshold.
New coverage¶
GhanaSPS gains four tables: sample (#673 -- data_scheme.yml had
asserted "cluster identity unavailable", which was false; 334 EAs in all three
waves), household_roster (#677 -- 54,251 people, and
household_characteristics with it), housing (#726) and
individual_education (#728). Each records the trap that would have made it
silently wrong: the module letter shifts between waves while the variable
names do not, so wave 3's 01fi_* is employment where wave 2's is education.
livestock is defined canonically, and its value column split by
additivity (#736). There was no canonical block, so three different
quantities were all spelled Value and all graded sane:
Feature('livestock') had been stacking 71,121 per-head prices (Uganda,
Nigeria, Malawi) on 23,514 sales totals (Ethiopia, Mali) in one column. Split
by what arithmetic is legal -- ValuePerAnimal (not additive), HerdValue (a
stock), SalesValue (a flow) -- the discriminating test being whether the value
is reported when the household sold nothing, settled against questionnaire
wording. Values are unchanged: rows, nulls, sums and describe() identical for
all 14 countries. tests/test_currency_livestock_registry.py pins
currency.py's second copy of "which columns hold money" against the canonical
YAML by set equality; a stale entry there has no symptom at all, and silently
leaves a renamed column unconverted.
The coverage matrix grades against a pinned feature vocabulary (#725,
closes #724). build_matrix took its feature axis from each country's own
config, so a feature a country never declared produced no row -- GhanaSPS
reported 21/21 = 100% sane while ~300 of its 328 source files sat unwired.
That pre-closed 417 (country, feature) pairs as "not applicable" with no
evidence and no record of who decided. The 22-feature vocabulary is measured,
not tuned: ranking the 33 source features by declaring-country count leaves a
real gap at 22 (nothing has 3, 4 or 5), so >=3 and >=6 select the identical
set. The shipped snapshot goes 1,849 -> 2,279 rows with 415 undeclared
cells; a not-asked or asked-not-distributed verdict closes one, so it is a
queue, not 415 permanently-red cells. It also fixes a pre-existing snapshot
bug: --no-readiness had been silently downgrading 1,217 sane cells to
declared in the committed snapshot.
labels= accepts a dict (#697, refs #682/#685) --
c.sample(labels={'Rural': 'Settlement'}) selects a label column for any
mapped column or index level. The scalar form still targets j only and is
byte-for-byte unchanged. Selection happens at the mapping site, keyed on
Code: a settlement ladder is coarse-to-fine, so keying on Preferred Label
gives two arbitrary entries for seven rungs. No country config is wired; this
is the mechanism. Cache hashes are unchanged, and a structural test walks the
real @build_transform closure so a refactor that drags the label path into
the build fingerprint fails a test rather than invalidating the corpus. Guide:
docs/guide/label-selection.md.
Tests, CI and documentation¶
--collect-onlyno longer purges the cache (#722, closes #719) -- collection is a dry run. pytest collection is scoped totests/(#733), after a permutation test shipped alongside a findings doc did its data loading at module level and aborted the whole CI session, naming neither a test nor the change under review. A guard checked withast, not text matching, catches a dependency shadowing thetestspackage (#687 carried this alongside the GLSS5 fix above; refs #680).docs/guide/population.mdanddocs/guide/label-selection.mdare new;coverage.mdgains theundeclaredtier andcaching.mdrecords weights as the clearest parquet-vs-API divergence.- Seventeen task ledgers under
.coder/ledger/, and 11 GhanaLSSabsent_verdictsrows closing 6 cells (#678) -- Ghana previously had zero. - GhanaSPS is not a post-planting/post-harvest country (#735, closes #730):
no wave ships two rounds. Its constructible agriculture features are sized as
todo(#729), andplotidis documented as not a panel id after wave 2 -- tested, negative. - A branch audit records why 26 #323 country fixes had never merged (#670).
Known open¶
- #323 and #637 remain open. The country fixes here are a large block of them, not the whole class.
- #709: Kyrgyz 1998
SECT_06.DTA:fprimary= -1.2255e294 is a distinct value pandas does not filter, and it is delivered. The pandas 3.0 floor closes the 2^333 exposure only. LSMS_READ_STRICT=1is not on in CI.optional:is country-grain while a null-read finding is wave-grain, so there is no way to record that Niger 2014-15's missing coordinates are correct without disarming the column for every wave. That gates the ratchet.- GLSS2's asserted weight of 1.0 assumes attrition from the GLSS1 rotating
panel was ignorable, which nobody has shown.
PANELC.DATcarries the linkage if a genuine panel weight is ever built. - #592: 18 waves across 6 countries still have zero GPS in
cluster_features. - EthiopiaRHS
livestockandincomeare misdated (#738). The seven{village}lvs5.taband four{village}inc5.tabfiles sit in the1999/wave directory and are served att='1999', but they are the 1989 IFPRI survey -- convergent evidence from SPSS_datestamps, the packed 1989HHIDkey, 95-98% household overlap with the 1989-wiredasset{v}.tabsiblings against zero overlap with the R5 roster, and a village set that is exactly 1989's seven PAs. 541 households are affected; 2004 and 2009 rows are not. Documented, deliberately not rewired -- re-dating a wave touchessample/cluster_features/panel_idsand the coverage matrix. This is the one country where the canonicallivestockschema above sits on top of data carrying the wrongt.
Upgrading¶
- pandas >= 3.0 is required. Installs pinned below it are excluded; see #710 for why.
- The first build after upgrading is a cold one. Nothing to do.
- If you use
sample().weight, check whether you relied on its scale. Weighted ratios, shares and means are unchanged to machine precision;sum(weight)is now a household count. If you need expansion weights, read the L2-country parquet, which keeps the raw values.