motherduck-build-data-pipeline
Installation
SKILL.md
Build a Data Pipeline with MotherDuck
Start Here: Is a MotherDuck Server Active?
Use an active remote MotherDuck MCP server or local MotherDuck server to inspect the in-scope database, schema, grain, keys, and relevant metrics. Reuse known context and narrow discovery to the requested work; do not scan the whole workspace by default. Let the actual data model shape the result.
Resolve the target from the request or active context. Ask only if ambiguity materially affects the result. Without a server, use supplied schema and explicit assumptions for planning; do not imply live validation.
Pipeline Defaults
- batch over streaming
- raw landing before curation
- explicit raw -> staging -> analytics boundaries
- bulk ingest paths over row-by-row writes
- idempotent stage rebuilds or append contracts before scheduled automation
- verify the MotherDuck-supported DuckDB client version before recommending upstream-only write, checkpoint, or lakehouse features
- native MotherDuck storage unless DuckLake is explicitly required
- MotherDuck CLI for Flight source and large file-shaped output when the agent has a shell; MCP for chat-only operation
- a
flightsGuide for reusable scheduling, naming, secret, and ingestion conventions when the organization has them