Programming Language Lineage Dataset
An open, evidence-backed dataset of programming language implementation and influence relationships. Every relationship includes a confidence score and at least one evidence source URL.
Current version: v5.0
Relationship Breakdown
| Type | Count |
|---|---|
| influenced | 252 |
| compiler written in | 96 |
| runtime written in | 67 |
| bootstrap written in | 15 |
| transpiled to | 11 |
| rewritten in | 2 |
Schema
Each language node contains: id, name, first_release_year, paradigm, typing, cluster_hint.
Each relationship contains: from_language, to_language, relationship, confidence (0–1), evidence_source (URL), notes.
Download
Download the complete v5.0 dataset as JSON. Counts and metadata on this page are generated from this file.
https://www.languagelineage.org/dataset/v5/lineage_v5.json
License
The Language Lineage dataset is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0). You may share and adapt the data for any purpose, provided you give appropriate credit and indicate whether changes were made.
Citation
Suggested citation:
Language Lineage. Programming Language Lineage Dataset, v5.0. 152 nodes and 443 relationships. Accessed 2026. https://www.languagelineage.org/dataset
Explore in Graph →