GitStar
TrendingExploreTop 100Insight
⌘K
GitStar

A ranking dashboard for GitHub momentum, durable repository leaders, package adoption, and editorial context.

Data source · GitHub API and package ecosystem snapshots

Home

HomeTrendingMomentumPulseExplore

Discover

CategoriesLanguagesOrganizationsAI / MLMCP

Workflow

CompareWatchlistRandom

Knowledge

InsightGuideMethodology

Support

FAQAbout GitStarContactPrivacyTerms

© 2026 GitStar. All rights reserved.

Data sourced from GitHub API

  1. Home
  2. Category
  3. Data
📊Category

Data is a mini landing page

Use this route to compare repositories that solve the same problem lane. The first read should tell you who leads the category, which language dominates it, and whether the surface is concentrated in a few repositories or spread across several active contenders.

Leader

Asabeneh/30-Days-Of-Python

55.7K stars

Top repository in the all time category snapshot.

Dominant language

Python

9 repos

Python (9), Unknown (4), TypeScript (3)

Shape insight

26% concentration

30 repos tracked

Top-three concentration shows whether this category is flagship-heavy or broad.

Recent activity coverage

0 recently active

All Time

Snapshot updated 2026-01-15.

Compare nearby reposRead methodology
1
Asabeneh/30-Days-Of-Python
55.7K stars·Python·10.7K forks·6 months ago

The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days. This challenge may take more than 100 days. Follow your own pace. These videos may help too: https://www.youtube.com/channel/UC7PNRuno1rzYPb1xLa4yktw

2
TanStack/query
48.2K stars·TypeScript·3.7K forks·6 months ago

🤖 Powerful asynchronous state management, server-state utilities and data fetching for the web. TS/JS, React Query, Solid Query, Svelte Query and Vue Query.

3
run-llama/llama_index
46.4K stars·Python·6.7K forks·6 months ago

LlamaIndex is the leading framework for building LLM-powered agents over your data.

4
metabase/metabase
45.6K stars·Clojure·6.2K forks·6 months ago

The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with data :bar_chart:

5
DataExpert-io/data-engineer-handbook
39.4K stars·Jupyter Notebook·7.6K forks·7 months ago

This is a repo with links to everything you'd ever want to learn about data engineering

6
SheetJS/sheetjs
36.1K stars·Language not captured·8K forks·2 years ago

📗 SheetJS Spreadsheet Data Toolkit -- New home https://git.sheetjs.com/SheetJS/sheetjs

7
vercel/swr
32.3K stars·TypeScript·1.3K forks·6 months ago

React Hooks for Data Fetching

8
sinaptik-ai/pandas-ai
23K stars·Python·2.3K forks·9 months ago

Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.

9
PrefectHQ/prefect
21.3K stars·Python·2.1K forks·6 months ago

Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

10
airbytehq/airbyte
20.5K stars·Python·5K forks·6 months ago

The leading data integration platform for ETL / ELT data pipelines from APIs, databases & files to data warehouses, data lakes & data lakehouses. Both self-hosted and Cloud-hosted.

11
fivethirtyeight/data
17.3K stars·Jupyter Notebook·11.1K forks·1 years ago

Data and code behind the articles and graphics at FiveThirtyEight

12
prestodb/presto
16.6K stars·Java·5.5K forks·6 months ago

The official home of the Presto distributed SQL query engine for big data

13
akfamily/akshare
15.4K stars·Python·2.7K forks·6 months ago

AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库

14
faker-js/faker
14.8K stars·TypeScript·1K forks·6 months ago

Generate massive amounts of fake data in the browser and node.js

15
oxnr/awesome-bigdata
14.2K stars·Language not captured·2.6K forks·8 months ago

A curated list of awesome big data frameworks, ressources and other awesomeness.

16
pwxcoo/chinese-xinhua
11.4K stars·Python·2.7K forks·2 years ago

:orange_book: 中华新华字典数据库。包括歇后语,成语,词语,汉字。

17
apple/pkl
11K stars·Java·352 forks·6 months ago

A configuration as code language with rich validation and tooling.

18
PRQL/prql
10.6K stars·Rust·250 forks·6 months ago

PRQL is a modern language for transforming data — a simple, powerful, pipelined SQL replacement

19
bchavez/Bogus
9.6K stars·C#·538 forks·7 months ago

:card_index: A simple fake data generator for C#, F#, and VB.NET. Based on and ported from the famed faker.js.

20
rawgraphs/rawgraphs-app
8.9K stars·JavaScript·1.9K forks·8 months ago

A web interface to create custom vector-based visualizations on top of RAWGraphs core

21
D4Vinci/Scrapling
8.8K stars·Python·534 forks·6 months ago

🕷️ An undetectable, powerful, flexible, high-performance Python library to make Web Scraping Easy and Effortless as it should be!

22
mage-ai/mage-ai
8.6K stars·Python·904 forks·6 months ago

🧙 Build, run, and manage data pipelines for integrating and transforming data.

23
mrdbourke/machine-learning-roadmap
7.8K stars·Language not captured·1.2K forks·3 years ago

A roadmap connecting many of the most important concepts in machine learning, how to learn them and what tools to use to perform them.

24
olifolkerd/tabulator
7.5K stars·JavaScript·877 forks·11 months ago

Interactive Tables and Data Grids for JavaScript

25
snowplow/snowplow
7K stars·Scala·1.2K forks·1 years ago

The leader in Customer Data Infrastructure

26
flyteorg/flyte
6.7K stars·Go·774 forks·6 months ago

Scalable and flexible workflow orchestration platform that seamlessly unifies data, ML and analytics stacks.

27
cloudquery/cloudquery
6.3K stars·Go·547 forks·6 months ago

Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.

28
dformoso/machine-learning-mindmap
6.3K stars·Language not captured·1K forks·6 years ago

A mindmap summarising Machine Learning concepts, from Data Analysis to Deep Learning.

29
axa-group/Parsr
6.2K stars·JavaScript·326 forks·2 years ago

Transforms PDF, Documents and Images into Enriched Structured Data

30
cue-lang/cue
5.9K stars·Go·347 forks·6 months ago

The home of the CUE language! Validate and define text-based and dynamic configuration

#RepositoryLanguageStars🍴 ForksUpdated
1
Asabeneh/30-Days-Of-Python

The 30 Days of Python programming challenge is a step-by-step guide to learn the Python programming language in 30 days. This challenge may take more than 100 days. Follow your own pace. These videos may help too: https://www.youtube.com/channel/UC7PNRuno1rzYPb1xLa4yktw

Python55.7K10.7K6 months ago
2
TanStack/query

🤖 Powerful asynchronous state management, server-state utilities and data fetching for the web. TS/JS, React Query, Solid Query, Svelte Query and Vue Query.

TypeScript48.2K3.7K6 months ago
3
run-llama/llama_index

LlamaIndex is the leading framework for building LLM-powered agents over your data.

Python46.4K6.7K6 months ago
4
metabase/metabase

The easy-to-use open source Business Intelligence and Embedded Analytics tool that lets everyone work with data :bar_chart:

Clojure45.6K6.2K6 months ago
5
DataExpert-io/data-engineer-handbook

This is a repo with links to everything you'd ever want to learn about data engineering

Jupyter Notebook39.4K7.6K7 months ago
6
SheetJS/sheetjs

📗 SheetJS Spreadsheet Data Toolkit -- New home https://git.sheetjs.com/SheetJS/sheetjs

36.1K8K2 years ago
7
vercel/swr

React Hooks for Data Fetching

TypeScript32.3K1.3K6 months ago
8
sinaptik-ai/pandas-ai

Chat with your database or your datalake (SQL, CSV, parquet). PandasAI makes data analysis conversational using LLMs and RAG.

Python23K2.3K9 months ago
9
PrefectHQ/prefect

Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

Python21.3K2.1K6 months ago
10
airbytehq/airbyte

The leading data integration platform for ETL / ELT data pipelines from APIs, databases & files to data warehouses, data lakes & data lakehouses. Both self-hosted and Cloud-hosted.

Python20.5K5K6 months ago
11
fivethirtyeight/data

Data and code behind the articles and graphics at FiveThirtyEight

Jupyter Notebook17.3K11.1K1 years ago
12
prestodb/presto

The official home of the Presto distributed SQL query engine for big data

Java16.6K5.5K6 months ago
13
akfamily/akshare

AKShare is an elegant and simple financial data interface library for Python, built for human beings! 开源财经数据接口库

Python15.4K2.7K6 months ago
14
faker-js/faker

Generate massive amounts of fake data in the browser and node.js

TypeScript14.8K1K6 months ago
15
oxnr/awesome-bigdata

A curated list of awesome big data frameworks, ressources and other awesomeness.

14.2K2.6K8 months ago
16
pwxcoo/chinese-xinhua

:orange_book: 中华新华字典数据库。包括歇后语,成语,词语,汉字。

Python11.4K2.7K2 years ago
17
apple/pkl

A configuration as code language with rich validation and tooling.

Java11K3526 months ago
18
PRQL/prql

PRQL is a modern language for transforming data — a simple, powerful, pipelined SQL replacement

Rust10.6K2506 months ago
19
bchavez/Bogus

:card_index: A simple fake data generator for C#, F#, and VB.NET. Based on and ported from the famed faker.js.

C#9.6K5387 months ago
20
rawgraphs/rawgraphs-app

A web interface to create custom vector-based visualizations on top of RAWGraphs core

JavaScript8.9K1.9K8 months ago
21
D4Vinci/Scrapling

🕷️ An undetectable, powerful, flexible, high-performance Python library to make Web Scraping Easy and Effortless as it should be!

Python8.8K5346 months ago
22
mage-ai/mage-ai

🧙 Build, run, and manage data pipelines for integrating and transforming data.

Python8.6K9046 months ago
23
mrdbourke/machine-learning-roadmap

A roadmap connecting many of the most important concepts in machine learning, how to learn them and what tools to use to perform them.

7.8K1.2K3 years ago
24
olifolkerd/tabulator

Interactive Tables and Data Grids for JavaScript

JavaScript7.5K87711 months ago
25
snowplow/snowplow

The leader in Customer Data Infrastructure

Scala7K1.2K1 years ago
26
flyteorg/flyte

Scalable and flexible workflow orchestration platform that seamlessly unifies data, ML and analytics stacks.

Go6.7K7746 months ago
27
cloudquery/cloudquery

Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.

Go6.3K5476 months ago
28
dformoso/machine-learning-mindmap

A mindmap summarising Machine Learning concepts, from Data Analysis to Deep Learning.

6.3K1K6 years ago
29
axa-group/Parsr

Transforms PDF, Documents and Images into Enriched Structured Data

JavaScript6.2K3262 years ago
30
cue-lang/cue

The home of the CUE language! Validate and define text-based and dynamic configuration

Go5.9K3476 months ago

Next step after the category scan

Open compare mode for a side-by-side read, move to the category hub, or review the guide when you need a stronger frame for what stars and activity are showing here.
Open compareBrowse all categoriesRead the guide

Learn and methodology

Keep trust-building context reachable, but behind the first data read instead of ahead of it.
GuideMethodologyArticlesWeekly Digest

How to use this category view

Category pages are most useful when you compare similar kinds of projects rather than treating the top row as an automatic recommendation list. Look at star concentration, language mix, and recent activity together to decide whether a category is dominated by a few established leaders or by a broader set of active tools.

Ranking methodology

Rankings use repositories tagged with the "data" topic, sorted by star count. We cover databases, query engines, visualization libraries, and pipeline orchestration tools.

Data is sourced from the GitHub REST API and updated daily. Star counts reflect cumulative community interest over the lifetime of each repository. See Methodology & Editorial Standards for broader context.