Country¶
The data methods are not on this page
Country.food_acquired(), .food_prices(), .household_roster() and every
other table method are generated at attribute-access time from the
country's data_scheme, so they have no source definition for this
reference to introspect. They — and the keyword arguments they share — are
documented in Data methods and their keyword arguments.
What follows is the introspection surface: how to ask a country what it has, not how to get the data out.
Country ¶
Country(country_name: str, preload_panel_ids: bool = False, verbose: bool = False, assume_cache_fresh: bool = False, trust_cache: bool = False)
Primary interface to a single country's LSMS survey data.
Provides access to all survey waves, standardized tables, and panel data.
Tables listed in data_scheme are available as callable attributes
(e.g. country.food_expenditures()).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
country_name
|
str
|
Directory name under |
required |
preload_panel_ids
|
bool
|
If True, compute panel ID mappings eagerly at construction time. Default is False (lazy). |
False
|
verbose
|
bool
|
Enable verbose logging. |
False
|
assume_cache_fresh
|
bool
|
If True, read existing cached Parquet files directly, bypassing DVC
and the normal build pipeline. Use this when you know the cache is
up-to-date and want to skip all existence / staleness checks.
.. note::
This is the one escape that also skips the v0.8.0 content-hash
staleness check — it is the explicit "I promise the cache
matches current sources" mode. With the hash check now the
default, |
False
|
trust_cache
|
bool
|
Deprecated alias for |
False
|
Examples:
>>> import lsms_library as ll
>>> uga = ll.Country('Uganda')
>>> uga.waves
['2005-06', '2009-10', ...]
>>> uga.data_scheme
['food_acquired', 'household_roster', ...]
>>> food = uga.food_expenditures()
categorical_mapping
property
¶
Get the categorical mapping for the country.
Searches current directory, then parent directory.
Also merges global .org files from lsms_library/categorical_mapping/ (GH #168).
Global tables are loaded first; per-country tables override on name
collision -- except names in _ADDITIVE_CATEGORICAL_TABLES (e.g.
u), which are row-unioned so a country inherits global rows and
overrides only the keys it redeclares (DESIGN_u_consolidation).
note_topics
property
¶
(level, TODO keyword, headline) for every heading in the notes.
The discovery half of :meth:notes, and the load-bearing half: the
headline vocabulary is largely ad hoc -- counts range from 11 headings
to 124 -- so nobody guesses 'Household Presence / MonthsSpent'.
This is to :meth:notes what :attr:data_scheme is to the table
methods.
population
property
¶
{wave: PopulationRecord} -- what each wave's sample REPRESENTS.
The universe is a property of the (country, wave) cell, not of the
country: Ethiopia ESS W1 covered rural areas and small towns while W4
is national, so a panel across them pools two populations. Read from
{country}/_/population.yml; see :mod:lsms_library.population.
Every record carries universe_tag (an EDITORIAL reading, never a
quote) together with source_type and confidence, which say where
the statement came from and how strongly it is a universe statement at
all. Quoting the tag without those two turns a judgement into a fact.
Empty for a country the 2026-07-21 sweep did not reach.
features
property
¶
The tables this country provides -- a synonym for :attr:data_scheme.
The library talks about features nearly everywhere: :class:Feature
assembles one across countries, the coverage matrix grades
(country, feature, wave) cells, and the guides are written in those
terms. The attribute that lists them was named instead after the
data_scheme.yml file it happens to read. Both names now work, so
the vocabulary a reader arrives with is the one that answers.
Deliberately a synonym and not a rename with a deprecation:
data_scheme is used throughout the countries' own scripts and in
published notebooks, and the file it is named for is not going away.
data_scheme
property
¶
List of data objects available for country.
Includes derived tables (e.g. food_expenditures, household_characteristics) when their source table is present, even if they are not explicitly registered in data_scheme.yml.
panel_ids
property
¶
Raw panel-ID tables keyed by wave. Computed lazily on first access.
Gated on _panel_ids_attempted rather than the cache value so a
legitimate negative result (country without panel design) is
cached as None without re-running _compute_panel_ids.
updated_ids
property
¶
Mapping {old_id: new_id} per wave for ID harmonization. Computed lazily.
notes ¶
This country's CONTENTS.org -- its recorded idiosyncrasies.
CONTENTS.org is where the repository records what is odd about a
survey: identifier conventions, design quirks, known defects, and
decisions already taken with their reasons. It was previously
reachable only by navigating the filesystem, which meant anyone
working through the API could not find it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
topic
|
str
|
Case-insensitive substring matched against headline text
(body text is not searched). Returns each matching headline with
its whole subtree, since the useful detail is nested -- a
country's |
None
|
state
|
str
|
TODO keyword to filter on, e.g. |
None
|
Returns:
| Type | Description |
|---|---|
str
|
The matching text, or |
See Also
note_topics : the headlines available to pass as topic.
derivations ¶
{key: DerivationRecord} -- served numbers that are NOT survey answers.
Every entry names a construction the library performs on this
country's data (a wave script, or a framework-level rule with an
empty country slot): the rule, the function that implements
it, the inputs callable that returns the raw answers by their
original names, and the assumptions with their basis. Rows
carrying a derivation are labelled in the served table's
Derivation column with the same key. Read from
{country}/_/derivations.yml and lsms_library/derivations.yml;
see :mod:lsms_library.derivations and
SkunkWorks/derived_values.org.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
table
|
str
|
Restrict to entries for this table (an entry with an empty table slot applies to every table). |
None
|
derivation_inputs ¶
The exact raw survey answers a derivation was computed from.
Resolves the entry's inputs callable and returns its frame: the
ORIGINAL variable names, at the input grain, re-read from the source
through get_dataframe. Slow and exact; nothing is cached. A user
who wants the value under a different assumption re-derives it from
this -- the derivation function itself takes no options.
The household level i is re-keyed through updated_ids exactly
as _finalize_result re-keys every served table, so the frame joins
to the served rows by (t, i, ...). Without this the raw frame
carries the survey's own household id -- GhanaLSS 1988-89 serves
'101332' / '101332_1' (a panel re-key plus a split-household
suffix, GH #548) where Y12B.DAT says '200103' -- and "the exact
inputs of this row" would not be findable from the row.
provenance ¶
Tabular survey of source + license per wave.
Reads <wave>/Documentation/SOURCE.org and
<wave>/Documentation/LICENSE.org for every wave in
:attr:waves. Useful for AER-style data-editor reviews where
the reviewer wants programmatic provenance verification
rather than poking at the filesystem.
Returns:
| Type | Description |
|---|---|
DataFrame
|
Indexed by wave label
|
Reads silently --- no warnings for waves missing one or both
|
|
files. Use :attr:`Wave.license` / :attr:`Wave.data_source`
|
|
directly if you want the warning side effect on miss.
|
|
cached_datasets ¶
List dataset names currently cached for this country.
Discovers caches at three locations under data_root()
(DVC blob cache L1 lives under dvc-cache/ and is not enumerated
by this method):
- L2-country:
data_root(country)/var/*.parquet. - Country-level companion:
data_root(country)/_/*.{parquet,json}. - L2-wave:
data_root(country)/{wave}/_/*.parquetfor every wave subdirectory (excluding the special_/andvar/directories above). This catches script-path tables (Nigeria's PP/PHhousehold_roster, Tanzania's multi-round tables) whose only cache lives at the wave level.
clear_cache ¶
clear_cache(methods: list[str] | None = None, waves: list[str] | None = None, dry_run: bool = False) -> list[Path]
Remove cached files for this country.
Clears two parquet tiers (the L1 DVC blob cache is left alone;
evict it manually with rm -rf {data_root}/dvc-cache if you
really mean to re-fetch from S3) plus DVC build artifacts:
- L2-country at
data_root(country)/var/{method}.parquetand the JSON-companiondata_root(country)/_/{method}.parquet|json. - L2-wave at
data_root(country)/{wave}/_/{method}.parquetfor every wave the country knows about (via :attr:waves). Both YAML-path tables (cached on firstWave.grab_datacall, since the 2026-04-15 L2-wave addition) and script-path tables (written by_/{method}.pyvia :func:local_tools.to_parquet) live here. - DVC build artifacts under
lsms_library/countries/{country}/...whenwaves=is passed (resolved via the materialize stage map).
If methods is None, all cached datasets are removed.
Returns list of deleted file paths.
panel_attrition ¶
Produce an upper-triangular matrix showing the number of households that transition between rounds.
Args: waves: List of waves to analyze (default: all waves) index: Index variable to track (default: 'i' for household) return_ids: If True, return both matrix and IDs dict split_households_new_sample: If True, treat split households as new sample
Returns: DataFrame showing panel structure, optionally with IDs dict
test_all_data_schemes ¶
Test whether all method_names in obj.data_scheme can be successfully built. Falls back to Makefile if not in data_scheme.