Skip to content

Country

The data methods are not on this page

Country.food_acquired(), .food_prices(), .household_roster() and every other table method are generated at attribute-access time from the country's data_scheme, so they have no source definition for this reference to introspect. They — and the keyword arguments they share — are documented in Data methods and their keyword arguments.

What follows is the introspection surface: how to ask a country what it has, not how to get the data out.

Country

Country(country_name: str, preload_panel_ids: bool = False, verbose: bool = False, assume_cache_fresh: bool = False, trust_cache: bool = False)

Primary interface to a single country's LSMS survey data.

Provides access to all survey waves, standardized tables, and panel data. Tables listed in data_scheme are available as callable attributes (e.g. country.food_expenditures()).

Parameters:

Name Type Description Default
country_name str

Directory name under lsms_library/countries/ (e.g. 'Uganda', 'Tanzania').

required
preload_panel_ids bool

If True, compute panel ID mappings eagerly at construction time. Default is False (lazy).

False
verbose bool

Enable verbose logging.

False
assume_cache_fresh bool

If True, read existing cached Parquet files directly, bypassing DVC and the normal build pipeline. Use this when you know the cache is up-to-date and want to skip all existence / staleness checks. _finalize_result (kinship expansion, canonical spelling, dtype coercion, _join_v_from_sample) still runs on every read — only the cache-lookup / DVC layer is bypassed. Useful on clusters where the parquet cache has been pre-built. Ignores LSMS_NO_CACHE.

.. note:: This is the one escape that also skips the v0.8.0 content-hash staleness check — it is the explicit "I promise the cache matches current sources" mode. With the hash check now the default, assume_cache_fresh is strictly weaker than the default path and is a candidate for deprecation (see SkunkWorks/dvc_object_management.org "Rethink trust_cache").

False
trust_cache bool

Deprecated alias for assume_cache_fresh. Will be removed in v0.8.0.

False

Examples:

>>> import lsms_library as ll
>>> uga = ll.Country('Uganda')
>>> uga.waves
['2005-06', '2009-10', ...]
>>> uga.data_scheme
['food_acquired', 'household_roster', ...]
>>> food = uga.food_expenditures()

categorical_mapping property

categorical_mapping: dict[str, DataFrame]

Get the categorical mapping for the country. Searches current directory, then parent directory. Also merges global .org files from lsms_library/categorical_mapping/ (GH #168). Global tables are loaded first; per-country tables override on name collision -- except names in _ADDITIVE_CATEGORICAL_TABLES (e.g. u), which are row-unioned so a country inherits global rows and overrides only the keys it redeclares (DESIGN_u_consolidation).

waves property

waves: list[str]

List of names of waves available for country.

notes_path property

notes_path: Path

Path to this country's _/CONTENTS.org.

note_topics property

note_topics: list[tuple[int, str | None, str]]

(level, TODO keyword, headline) for every heading in the notes.

The discovery half of :meth:notes, and the load-bearing half: the headline vocabulary is largely ad hoc -- counts range from 11 headings to 124 -- so nobody guesses 'Household Presence / MonthsSpent'. This is to :meth:notes what :attr:data_scheme is to the table methods.

population property

population: dict[str, 'PopulationRecord']

{wave: PopulationRecord} -- what each wave's sample REPRESENTS.

The universe is a property of the (country, wave) cell, not of the country: Ethiopia ESS W1 covered rural areas and small towns while W4 is national, so a panel across them pools two populations. Read from {country}/_/population.yml; see :mod:lsms_library.population.

Every record carries universe_tag (an EDITORIAL reading, never a quote) together with source_type and confidence, which say where the statement came from and how strongly it is a universe statement at all. Quoting the tag without those two turns a judgement into a fact.

Empty for a country the 2026-07-21 sweep did not reach.

features property

features: list[str]

The tables this country provides -- a synonym for :attr:data_scheme.

The library talks about features nearly everywhere: :class:Feature assembles one across countries, the coverage matrix grades (country, feature, wave) cells, and the guides are written in those terms. The attribute that lists them was named instead after the data_scheme.yml file it happens to read. Both names now work, so the vocabulary a reader arrives with is the one that answers.

Deliberately a synonym and not a rename with a deprecation: data_scheme is used throughout the countries' own scripts and in published notebooks, and the file it is named for is not going away.

data_scheme property

data_scheme: list[str]

List of data objects available for country.

Includes derived tables (e.g. food_expenditures, household_characteristics) when their source table is present, even if they are not explicitly registered in data_scheme.yml.

panel_ids property

panel_ids: dict[str, Any] | None

Raw panel-ID tables keyed by wave. Computed lazily on first access.

Gated on _panel_ids_attempted rather than the cache value so a legitimate negative result (country without panel design) is cached as None without re-running _compute_panel_ids.

updated_ids property

updated_ids: dict[str, dict[str, str]] | None

Mapping {old_id: new_id} per wave for ID harmonization. Computed lazily.

notes

notes(topic: str | None = None, state: str | None = None) -> str

This country's CONTENTS.org -- its recorded idiosyncrasies.

CONTENTS.org is where the repository records what is odd about a survey: identifier conventions, design quirks, known defects, and decisions already taken with their reasons. It was previously reachable only by navigating the filesystem, which meant anyone working through the API could not find it.

Parameters:

Name Type Description Default
topic str

Case-insensitive substring matched against headline text (body text is not searched). Returns each matching headline with its whole subtree, since the useful detail is nested -- a country's Weights and Strata live under its Sampling Design. When a parent and a descendant both match, only the parent is returned; it already contains the descendant.

None
state str

TODO keyword to filter on, e.g. 'WAITING'. A closed GitHub issue can still leave a live caveat parked here, so notes(state='WAITING') is the quick answer to "what is still open for this country?".

None

Returns:

Type Description
str

The matching text, or '' when nothing matches. A country with no CONTENTS.org warns and returns '', following :meth:Wave.license.

See Also

note_topics : the headlines available to pass as topic.

derivations

derivations(table: str | None = None) -> dict

{key: DerivationRecord} -- served numbers that are NOT survey answers.

Every entry names a construction the library performs on this country's data (a wave script, or a framework-level rule with an empty country slot): the rule, the function that implements it, the inputs callable that returns the raw answers by their original names, and the assumptions with their basis. Rows carrying a derivation are labelled in the served table's Derivation column with the same key. Read from {country}/_/derivations.yml and lsms_library/derivations.yml; see :mod:lsms_library.derivations and SkunkWorks/derived_values.org.

Parameters:

Name Type Description Default
table str

Restrict to entries for this table (an entry with an empty table slot applies to every table).

None

derivation_inputs

derivation_inputs(key: str, wave: str | None = None) -> pd.DataFrame

The exact raw survey answers a derivation was computed from.

Resolves the entry's inputs callable and returns its frame: the ORIGINAL variable names, at the input grain, re-read from the source through get_dataframe. Slow and exact; nothing is cached. A user who wants the value under a different assumption re-derives it from this -- the derivation function itself takes no options.

The household level i is re-keyed through updated_ids exactly as _finalize_result re-keys every served table, so the frame joins to the served rows by (t, i, ...). Without this the raw frame carries the survey's own household id -- GhanaLSS 1988-89 serves '101332' / '101332_1' (a panel re-key plus a split-household suffix, GH #548) where Y12B.DAT says '200103' -- and "the exact inputs of this row" would not be findable from the row.

provenance

provenance() -> pd.DataFrame

Tabular survey of source + license per wave.

Reads <wave>/Documentation/SOURCE.org and <wave>/Documentation/LICENSE.org for every wave in :attr:waves. Useful for AER-style data-editor reviews where the reviewer wants programmatic provenance verification rather than poking at the filesystem.

Returns:

Type Description
DataFrame

Indexed by wave label t. Columns:

  • source — verbatim contents of SOURCE.org, or pd.NA if the file is missing. Typically a World Bank Microdata catalog URL.
  • license — verbatim contents of LICENSE.org, or pd.NA. May be a URL, full terms text, or both, depending on what's recorded for that wave.
  • documentation_path — filesystem path to the wave's Documentation/ directory.
Reads silently --- no warnings for waves missing one or both
files. Use :attr:`Wave.license` / :attr:`Wave.data_source`
directly if you want the warning side effect on miss.

cached_datasets

cached_datasets() -> list[str]

List dataset names currently cached for this country.

Discovers caches at three locations under data_root() (DVC blob cache L1 lives under dvc-cache/ and is not enumerated by this method):

  • L2-country: data_root(country)/var/*.parquet.
  • Country-level companion: data_root(country)/_/*.{parquet,json}.
  • L2-wave: data_root(country)/{wave}/_/*.parquet for every wave subdirectory (excluding the special _/ and var/ directories above). This catches script-path tables (Nigeria's PP/PH household_roster, Tanzania's multi-round tables) whose only cache lives at the wave level.

clear_cache

clear_cache(methods: list[str] | None = None, waves: list[str] | None = None, dry_run: bool = False) -> list[Path]

Remove cached files for this country.

Clears two parquet tiers (the L1 DVC blob cache is left alone; evict it manually with rm -rf {data_root}/dvc-cache if you really mean to re-fetch from S3) plus DVC build artifacts:

  • L2-country at data_root(country)/var/{method}.parquet and the JSON-companion data_root(country)/_/{method}.parquet|json.
  • L2-wave at data_root(country)/{wave}/_/{method}.parquet for every wave the country knows about (via :attr:waves). Both YAML-path tables (cached on first Wave.grab_data call, since the 2026-04-15 L2-wave addition) and script-path tables (written by _/{method}.py via :func:local_tools.to_parquet) live here.
  • DVC build artifacts under lsms_library/countries/{country}/... when waves= is passed (resolved via the materialize stage map).

If methods is None, all cached datasets are removed. Returns list of deleted file paths.

panel_attrition

panel_attrition(waves=None, index='i', return_ids=False, split_households_new_sample=True)

Produce an upper-triangular matrix showing the number of households that transition between rounds.

Args: waves: List of waves to analyze (default: all waves) index: Index variable to track (default: 'i' for household) return_ids: If True, return both matrix and IDs dict split_households_new_sample: If True, treat split households as new sample

Returns: DataFrame showing panel structure, optionally with IDs dict

test_all_data_schemes

test_all_data_schemes(waves: list[str] | None = None) -> dict[str, str]

Test whether all method_names in obj.data_scheme can be successfully built. Falls back to Makefile if not in data_scheme.