GitStar
⌘K
GitStar

A ranking dashboard for GitHub momentum, durable repository leaders, package adoption, and editorial context.

Data source · GitHub API and package ecosystem snapshots

Home

HomeTrendingMomentumPulseExplore

Discover

CategoriesLanguagesOrganizationsAI / MLMCP

Workflow

CompareWatchlistRandom

Knowledge

InsightGuideMethodology

Support

FAQAbout GitStarContactPrivacyTerms

© 2026 GitStar. All rights reserved.

Data sourced from GitHub API

TrendingMomentumPulseExplore
  1. Home
  2. Category
  3. Data
📊CATEGORY LEADERBOARD

Data

Data processing and analytics

Repositories
96
Combined Stars
3.2M
Top-3 Concentration
9%
Recently Active
73Updated Sep 12, 2026
Language Distribution
Python36.0%
Unknown13.0%
Rust9.0%
TypeScript8.0%
Jupyter Notebook7.0%
C++7.0%
Topics

Top Repositories

Top 3
1st#1

supabase/supabase

The Postgres development platform. Supabase gives you a dedicated Postgres database to build your web, mobile, and AI applications.

109.1KStars
13.7KForks
YesterdayUpdated
TypeScriptApache-2.0since 2019
2nd#2

microsoft/ML-For-Beginners

12 weeks, 26 lessons, 52 quizzes, classic Machine Learning for all

90.4KStars
22.3KForks
YesterdayUpdated
Jupyter NotebookMITsince 2021
3rd#3

netdata/netdata

The fastest path to AI-powered full stack observability, even for lean teams.

80.5KStars
6.6KForks
YesterdayUpdated
GoGPL-3.0since 2013

Full ranking

30
#
Repository
Language
Stars
Forks
Updated
04
redis/redis
For developers, who are building real-time data-driven applications, Redis is the preferred, fastest, and most feature-rich cache, data structure server, and document and vector query engine.
est. 2009
C
76.3K
24.8K
3 days ago
05
apache/superset
Apache Superset is a Data Visualization and Data Exploration Platform
est. 2015
Python
74.7K
18.3K
Yesterday
06
Asabeneh/30-Days-Of-Python
The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days. This challenge may take more than 100 days. Follow your own pace. These videos may help too: https://www.youtube.com/channel/UC7PNRuno1rzYPb1xLa4yktw
est. 2019
Python
73.6K
13.4K
3 days ago
07
scikit-learn/scikit-learn
scikit-learn: machine learning in Python
est. 2010
Python
67.2K
27.4K
Yesterday
08
keras-team/keras
Deep Learning for humans
est. 2015
Python
64.3K
19.8K
Yesterday
09
meilisearch/meilisearch
A lightning-fast search engine API bringing AI-powered hybrid search to your sites and applications.
est. 2018
Rust
59.3K
2.7K
3 days ago
10
etcd-io/etcd
Distributed reliable key-value store for the most critical data of a distributed system
est. 2013
Go
52.3K
10.5K
2 days ago
11
dbeaver/dbeaver
Free universal database tool and SQL client
est. 2015
Java
51.7K
4.4K
Yesterday
12
ClickHouse/ClickHouse
ClickHouse® is a real-time analytics database management system
est. 2016
C++
49.8K
9K
Yesterday
13
pandas-dev/pandas
Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more
est. 2010
Python
49.7K
20.4K
Yesterday
14
GokuMohandas/Made-With-ML
Learn how to develop, deploy and iterate on production-grade ML applications.
est. 2018
Jupyter Notebook
49.5K
7.8K
6 months ago
15
metabase/metabase
The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with data :bar_chart:
est. 2015
Clojure
49.2K
6.8K
Yesterday
16
prisma/orm
Next-generation ORM for Node.js & TypeScript | PostgreSQL, MySQL, MariaDB, SQL Server, SQLite, MongoDB and CockroachDB
est. 2019
TypeScript
47.6K
2.5K
Yesterday
17
SimplifyJobs/Summer2027-Internships
Summer 2027 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.
est. 2020
Python
47.4K
3.2K
Yesterday
18
apache/airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
est. 2015
Python
46.8K
17.8K
Yesterday
19
LeCoupa/awesome-cheatsheets
👩‍💻👨‍💻 Awesome cheatsheets for popular programming languages, frameworks and development tools. They include everything you should know in one single file.
est. 2017
JavaScript
46.5K
6.7K
5 months ago
20
streamlit/streamlit
Streamlit — A faster way to build and share data apps.
est. 2019
Python
45.7K
4.4K
Yesterday
21
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
est. 2016
Python
43.8K
8K
Yesterday
22
gradio-app/gradio
Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!
est. 2018
Python
43.5K
3.6K
Yesterday
23
duckdb/duckdb
DuckDB is an analytical in-process SQL database management system
est. 2018
C++
41.2K
3.8K
2 days ago
24
pingcap/tidb
TiDB is built for agentic workloads that grow unpredictably, with ACID guarantees and native support for transactions, analytics, and vector search. No data silos. No noisy neighbors. No infrastructure ceiling.
est. 2015
Go
40.5K
6.2K
Yesterday
25
The-Vibe-Company/quivr
Opiniated RAG for integrating GenAI in your apps 🧠 Focus on your product rather than the RAG. Easy integration in existing products with customisation! Any LLM: GPT4, Groq, Llama. Any Vectorstore: PGVector, Faiss. Any Files. Anyway you want.
est. 2023
Python
39.5K
3.7K
1 weeks ago
26
huihut/interview
📚 C/C++ 技术面试基础知识总结,包括语言、程序库、数据结构、算法、系统、网络、链接装载库等知识及面试经验、招聘、内推等信息。This repository is a summary of the basic knowledge of recruiting job seekers and beginners in the direction of C/C++ technology, including language, program library, data structure, algorithm, system, network, link loading library, interview experience, recruitment, recommendation, etc.
est. 2018
C++
38.2K
8.1K
1 years ago
27
directus/directus
The flexible backend for all your projects 🐰 Turn your DB into a headless CMS, admin panels, or apps with a custom UI, instant APIs, auth & more.
est. 2012
TypeScript
37.9K
4.9K
2 days ago
28
microsoft/Data-Science-For-Beginners
10 Weeks, 20 Lessons, Data Science for All!
est. 2021
Jupyter Notebook
36.9K
7.5K
Yesterday
29
ashishpatel26/500-AI-Machine-learning-Deep-learning-Computer-vision-NLP-Projects-with-code
500 AI Machine learning Deep learning Computer vision NLP Projects with code
est. 2021
Unknown
36.8K
7.5K
1 years ago
30
typeorm/typeorm
TypeScript & JavaScript ORM for Node.js — supports PostgreSQL, MySQL, MariaDB, SQLite, SQL Server, Oracle, and more.
est. 2016
TypeScript
36.7K
6.7K
2 days ago
04
redis/redis
For developers, who are building real-time data-driven applications, Redis is the preferred, fastest, and most feature-rich cache, data structure server, and document and vector query engine.
C
★76.3K
🍴24.8K
est. 2009
05
apache/superset
Apache Superset is a Data Visualization and Data Exploration Platform
Python
★74.7K
🍴18.3K
est. 2015
06
Asabeneh/30-Days-Of-Python
The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days. This challenge may take more than 100 days. Follow your own pace. These videos may help too: https://www.youtube.com/channel/UC7PNRuno1rzYPb1xLa4yktw
Python
★73.6K
🍴13.4K
est. 2019
07
scikit-learn/scikit-learn
scikit-learn: machine learning in Python
Python
★67.2K
🍴27.4K
est. 2010
08
keras-team/keras
Deep Learning for humans
Python
★64.3K
🍴19.8K
est. 2015
09
meilisearch/meilisearch
A lightning-fast search engine API bringing AI-powered hybrid search to your sites and applications.
Rust
★59.3K
🍴2.7K
est. 2018
10
etcd-io/etcd
Distributed reliable key-value store for the most critical data of a distributed system
Go
★52.3K
🍴10.5K
est. 2013
11
dbeaver/dbeaver
Free universal database tool and SQL client
Java
★51.7K
🍴4.4K
est. 2015
12
ClickHouse/ClickHouse
ClickHouse® is a real-time analytics database management system
C++
★49.8K
🍴9K
est. 2016
13
pandas-dev/pandas
Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more
Python
★49.7K
🍴20.4K
est. 2010
14
GokuMohandas/Made-With-ML
Learn how to develop, deploy and iterate on production-grade ML applications.
Jupyter Notebook
★49.5K
🍴7.8K
est. 2018
15
metabase/metabase
The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with data :bar_chart:
Clojure
★49.2K
🍴6.8K
est. 2015
16
prisma/orm
Next-generation ORM for Node.js & TypeScript | PostgreSQL, MySQL, MariaDB, SQL Server, SQLite, MongoDB and CockroachDB
TypeScript
★47.6K
🍴2.5K
est. 2019
17
SimplifyJobs/Summer2027-Internships
Summer 2027 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.
Python
★47.4K
🍴3.2K
est. 2020
18
apache/airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
Python
★46.8K
🍴17.8K
est. 2015
19
LeCoupa/awesome-cheatsheets
👩‍💻👨‍💻 Awesome cheatsheets for popular programming languages, frameworks and development tools. They include everything you should know in one single file.
JavaScript
★46.5K
🍴6.7K
est. 2017
20
streamlit/streamlit
Streamlit — A faster way to build and share data apps.
Python
★45.7K
🍴4.4K
est. 2019
21
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Python
★43.8K
🍴8K
est. 2016
22
gradio-app/gradio
Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!
Python
★43.5K
🍴3.6K
est. 2018
23
duckdb/duckdb
DuckDB is an analytical in-process SQL database management system
C++
★41.2K
🍴3.8K
est. 2018
24
pingcap/tidb
TiDB is built for agentic workloads that grow unpredictably, with ACID guarantees and native support for transactions, analytics, and vector search. No data silos. No noisy neighbors. No infrastructure ceiling.
Go
★40.5K
🍴6.2K
est. 2015
25
The-Vibe-Company/quivr
Opiniated RAG for integrating GenAI in your apps 🧠 Focus on your product rather than the RAG. Easy integration in existing products with customisation! Any LLM: GPT4, Groq, Llama. Any Vectorstore: PGVector, Faiss. Any Files. Anyway you want.
Python
★39.5K
🍴3.7K
est. 2023
26
huihut/interview
📚 C/C++ 技术面试基础知识总结,包括语言、程序库、数据结构、算法、系统、网络、链接装载库等知识及面试经验、招聘、内推等信息。This repository is a summary of the basic knowledge of recruiting job seekers and beginners in the direction of C/C++ technology, including language, program library, data structure, algorithm, system, network, link loading library, interview experience, recruitment, recommendation, etc.
C++
★38.2K
🍴8.1K
est. 2018
27
directus/directus
The flexible backend for all your projects 🐰 Turn your DB into a headless CMS, admin panels, or apps with a custom UI, instant APIs, auth & more.
TypeScript
★37.9K
🍴4.9K
est. 2012
28
microsoft/Data-Science-For-Beginners
10 Weeks, 20 Lessons, Data Science for All!
Jupyter Notebook
★36.9K
🍴7.5K
est. 2021
29
ashishpatel26/500-AI-Machine-learning-Deep-learning-Computer-vision-NLP-Projects-with-code
500 AI Machine learning Deep learning Computer vision NLP Projects with code
Unknown
★36.8K
🍴7.5K
est. 2021
30
typeorm/typeorm
TypeScript & JavaScript ORM for Node.js — supports PostgreSQL, MySQL, MariaDB, SQLite, SQL Server, Oracle, and more.
TypeScript
★36.7K
🍴6.7K
est. 2016
Data via public GitHub API · refreshed every 8h30 of 96 shown

Next step after the category scan

Open compare mode for a side-by-side read, move to the category hub, or review the guide when you need a stronger frame for what stars and activity are showing here.
Open compareBrowse all categoriesRead the guide

Learn and methodology

Keep trust-building context reachable, but behind the first data read instead of ahead of it.
GuideMethodologyArticlesWeekly Digest

Data rankings: how to read the landscape behind the stars

The Data category tracks tools for data processing, database management, analytics pipelines, and ETL workflows. From Apache Spark to DuckDB, these projects handle the full spectrum of data engineering challenges.

Data is the fuel of modern business, and the tools used to collect, transform, and analyze it directly impact decision-making quality. The highest-starred data projects typically offer the best balance of performance, scalability, and developer experience.

How to use this category view

Category pages are most useful when you compare similar kinds of projects rather than treating the top row as an automatic recommendation list. Look at star concentration, language mix, and recent activity together to decide whether a category is dominated by a few established leaders or by a broader set of active tools.

Ranking methodology

Rankings use repositories tagged with the "data" topic, sorted by star count. We cover databases, query engines, visualization libraries, and pipeline orchestration tools.

Data is sourced from the GitHub REST API and updated daily. Star counts reflect cumulative community interest over the lifetime of each repository. See Methodology & Editorial Standards for broader context.