# dlt Source: https://www.hotdata.dev/docs/dlthub Site index: https://www.hotdata.dev/llms.txt Use [dlt](https://dlthub.com) to build pipelines that load data from any source into Hotdata instant databases — with automatic schema inference, incremental loading, and Parquet-based delivery. The Hotdata destination is a native dlt destination published as a standalone package, [`hotdata-dlt-destination`](https://github.com/hotdata-dev/hotdata-dlt-destination). ## Install ```bash pip install hotdata-dlt-destination ``` For the live ibis backend (`pipeline.dataset().ibis()`), add the `ibis` extra: ```bash pip install "hotdata-dlt-destination[ibis]" ``` ## Authentication The API key is a secret, so it's read from the environment (or a dlt secrets provider): ```bash export HOTDATA_API_KEY="your_api_key" ``` The workspace ID is a routing value, not a secret — pass it as the `workspace_id=` parameter on the destination (there is no environment variable for it). You can also supply the key explicitly with `credentials={"api_key": "..."}` instead of the env var. ## Quickstart ```python import dlt from hotdata_dlt_destination import hotdata @dlt.resource(name="customers", write_disposition="append") def customers(): yield [ {"id": 1, "name": "Alice", "amount": 99.99}, {"id": 2, "name": "Bob", "amount": 49.50}, ] pipeline = dlt.pipeline( pipeline_name="my_pipeline", destination=hotdata( workspace_id="your_workspace_id", database_name="sales", declared_tables=["customers"], ), ) info = pipeline.run(customers()) print(info) ``` dlt infers the schema from your data, creates an instant database labelled `sales` on the first run, and loads the records as Parquet. Nested/child tables and dlt's internal columns (`_dlt_id`, `_dlt_load_id`) are preserved. ## Reuse a database with `database_id` Instant databases are addressed by **id**, not name — Hotdata database names are not unique, so a name can't identify one. On first run the destination prints the new database's id: ``` hotdata: created managed database dbid6lguax1dxn9y1xj5gxnameyywl (name='sales'). Pin it for future runs by setting database_id=dbid6lguax1dxn9y1xj5gxnameyywl. ``` To keep loading into the **same** database on later runs, pin that id — via `hotdata(database_id="dbid6lguax1dxn9y1xj5gxnameyywl")`, the `HOTDATA_DATABASE_ID` environment variable, or `[destination.hotdata] database_id` in `.dlt/config.toml`. **Without a pinned id, every run creates a new database.** `database_name` is only a display label used when creating a database; it never looks one up. ## Configure the destination ```python from hotdata_dlt_destination import hotdata destination = hotdata( workspace_id="your_workspace_id", # required (no env var) # reuse an existing database (recommended) database_id="dbid6lguax1dxn9y1xj5gxnameyywl", database_name="sales", # label when creating a new database (default: "dlt") schema="public", # schema within the database (default: "public") write_disposition="append", # default disposition (default: "append") declared_tables=["customers"], # all table names the pipeline writes create_database_if_missing=True, # auto-create the database (default: True) ) ``` The API key is the exception — being a secret it comes from `HOTDATA_API_KEY` (or `credentials={"api_key": "..."}`), not a keyword above. Every other parameter can also be set via environment variable or `.dlt/config.toml`: | Parameter | Env variable | Default | |-----------|--------------|---------| | `workspace_id` | — (param only) | required | | `database_id` | `HOTDATA_DATABASE_ID` | — | | `database_name` | `HOTDATA_DATABASE` | `dlt` | | `schema` | `HOTDATA_SCHEMA` | `public` | | `write_disposition` | `HOTDATA_WRITE_DISPOSITION` | `append` | | `declared_tables` | `HOTDATA_DECLARED_TABLES` | — | | `create_database_if_missing` | `HOTDATA_CREATE_DATABASE_IF_MISSING` | `True` | | `api_base_url` | `HOTDATA_API_BASE_URL` | `https://api.hotdata.dev` | When a pipeline writes more than one table, pass every table name via `declared_tables`. If you add a new table later, include it on the next run — it's added to the existing database in place, without recreating it or moving data. ## Load from a source Use any dlt-verified source or a custom generator: ```python import dlt from dlt.sources.sql_database import sql_database from hotdata_dlt_destination import hotdata source = sql_database( credentials="postgresql://user:pass@host/db", schema="public", table_names=["orders", "customers"], ) pipeline = dlt.pipeline( pipeline_name="postgres_to_hotdata", destination=hotdata( workspace_id="your_workspace_id", database_name="sales", declared_tables=["orders", "customers"], ), ) info = pipeline.run(source) print(f"Loaded {info.loads_ids} into Hotdata") ``` ## Incremental loading dlt tracks state between runs — pipeline state is persisted in the instant database, so only new or updated rows are loaded on subsequent executions. Pin `database_id` so the state is found on the next run: ```python import dlt from hotdata_dlt_destination import hotdata @dlt.resource(primary_key="id", write_disposition="merge") def events( updated_at=dlt.sources.incremental("updated_at") ): # fetch rows newer than updated_at.last_value yield fetch_events(since=updated_at.last_value) pipeline = dlt.pipeline( pipeline_name="events_pipeline", destination=hotdata( workspace_id="your_workspace_id", database_id="dbid6lguax1dxn9y1xj5gxnameyywl", declared_tables=["events"], ), ) pipeline.run(events()) ``` Write dispositions: | Disposition | Behaviour | |-------------|-----------| | `append` | Add new rows to the table | | `replace` | Replace the full table on each run | | `merge` | Upsert rows matched by `primary_key` (updates matches, inserts the rest); without a `primary_key` it falls back to a client-side combine | A table's key is declared the first time it's created — changing a resource's `primary_key` on a later run does not update the server-side key. ## Partition and sort keys A managed table's partition and sort keys are fixed when the table is first created, so declare them before the first load. Simple cases use per-column hints: ```python @dlt.resource(columns={ "event_date": {"partition": True}, "event_time": {"sort": True}, }) def events(): ... ``` For key order, partition transforms (`year` / `month` / `day` / `hour`), or sort direction, use the adapter: ```python from hotdata_dlt_destination import hotdata_adapter hotdata_adapter( events, partition_by=[("event_date", "identity")], sorted_by=["event_time", ("tag_mac", "asc", "last")], ) ``` ## Read your data back The same dlt [dataset interface](https://dlthub.com/docs/general-usage/dataset-access/dataset) reads loaded tables back — queries run server-side on Hotdata's Apache DataFusion engine. Point the reading pipeline at the `database_id` you pinned: ```python pipeline = dlt.pipeline( pipeline_name="my_pipeline", destination=hotdata( workspace_id="your_workspace_id", database_id="dbid6lguax1dxn9y1xj5gxnameyywl", declared_tables=["customers"], ), ) ds = pipeline.dataset() ds.table("customers").df() # whole table -> pandas.DataFrame ds.table("customers").arrow() # -> pyarrow.Table # raw SQL ds("SELECT name, sum(amount) AS spend FROM customers GROUP BY name").df() # fluent ds.table("customers").select("id", "amount").where("amount > 50").limit(10).df() ``` With the `[ibis]` extra, `ds.ibis()` returns a live [`ibis.hotdata` backend](/docs/ibis) for authoring queries as ibis expressions. ## Verify a load Use the [Hotdata CLI](/docs/quick-start) to confirm the data landed. Address the database by the id printed on first-run create — names aren't unique: ```bash # List instant databases (shows each id) hotdata databases list # Query the loaded data hotdata query "SELECT name, amount FROM public.customers ORDER BY amount DESC" --database dbid6lguax1dxn9y1xj5gxnameyywl ``` ## See also - [hotdata-dlt-destination on GitHub](https://github.com/hotdata-dev/hotdata-dlt-destination) - [dlt documentation](https://dlthub.com/docs) - [Python SDK](/docs/python-sdk) — low-level `hotdata` HTTP API client - [Quick Start](/docs/quick-start) — CLI and workspace setup