Data Lineage Designer
Document where a field comes from and where its value is used. Create source, transformation and consumer nodes, connect explicit field dependencies, and trace downstream impact or upstream origins. This is a manual design workspace; it does not connect to databases or execute SQL.
Key features
- Separate node and field editors with source, transform and consumer classifications
- Explicit fan-in and fan-out mappings with notes for each dependency
- Downstream and upstream field reachability with shortest mapping counts
- Field-level cycle groups, isolated-field counts and complete reference tables
- Strict project JSON round-trips, a portable SVG drawing and impact CSV export
How to use
- Open the example or enter a node name and classification to add a node.
- Add fields to the selected node, then save mappings with source, destination and notes.
- Choose a starting field and run downstream impact or upstream origins.
- Review reached fields, affected nodes, shortest distances and cycle groups in the table and graph.
- Save project JSON for later editing, or export the current SVG and impact CSV.
Use cases
- Review revenue summaries and dashboards before changing the raw order amount field
- Trace a reporting metric back to its explicitly recorded input fields
- Explain merge and branching dependencies during an ETL migration review
- Share a manually documented lineage model as JSON and reopen it for editing
Frequently asked questions
How does this differ from a SQL formatter or ER diagram?
It does not format queries or define table-key relationships. It follows field dependencies that you explicitly record to calculate origins and potential impact. That also differs from the existing SQL formatter’s FROM/JOIN table reference summary.
Does a change affect every field in the same transformation node?
No. Only explicit field connections propagate impact. In the example, the raw amount reaches cleaned amount, revenue summary and dashboard revenue, while customer_id remains independent even when it belongs to the same node. Add every dependency you want considered.
Can I analyze a project that contains cycles?
Yes. Each field is visited once during reachability analysis. Actual field-level cycle groups and self-loops are listed separately. Opposite node-level arrows between unrelated fields do not automatically form a field cycle. Source, transform and consumer kinds are display classifications, not restrictions on direction.
What project JSON is accepted?
The local-data-lineage-v1 format requires exact properties, valid IDs and existing field endpoints. Duplicate decoded JSON properties, unknown properties, duplicate IDs, dangling endpoints and duplicate directed mappings are rejected. Limits are 256KiB, 40 nodes, 20 fields per node, 400 fields in total and 800 mappings.
What do the SVG and CSV contain?
SVG draws the saved nodes, fields and mappings and highlights the starting and reached fields after analysis. Long diagram labels are shortened; the project and reference tables retain their full text. CSV lists reached fields, excluding the starting field, with IDs, names, shortest distances and the first mapping that reached them. Formula-like strings receive an apostrophe prefix.
Does this discover real lineage or modify a pipeline?
No database connection, SQL inference, runtime-log collection or pipeline modification occurs. Results describe potential dependencies in your declared connections. Missing relationships and the actual execution of conditional transformations are not verified.
Privacy
Titles, fields, mappings, notes and files stay in this page’s memory. This tool does not upload them or automatically save them to a URL or browser storage. Download project JSON before reloading or leaving. A native file read cannot be stopped at the operating-system stage, but cancellation and newer input prevent its late result from being applied.
Comments & questions