Looking for a good example dataset for panel data analysis – any recommendations? or Where can I find

14 Replies, 1053 Views

Hey everyone,

I'm trying to get some hands-on practice with panel data analysis, but I'm struggling to find a solid example dataset for panel data analysis to work with.

Anyone got recommendations? Something clean but not *too* simple—maybe with a mix of time and cross-sectional dimensions?

I've seen a few options online, but not sure which ones are actually useful for learning. If you've used one before, lmk what worked for you!

Also, if there's a dataset that’s kinda fun or relatable (like household spending or student performance over time), that’d be a plus.

Thanks in advance!
Hey! You should check out the World Bank's open data portal. They have tons of country-level panel datasets with both time and cross-sectional dimensions.

For something more relatable, the "National Longitudinal Survey of Youth" (NLSY) is great—it tracks individuals over time, so you can analyze stuff like education, income, etc.

If you want something smaller, the "Penn World Table" is clean and widely used for macro panel data analysis.

Hope that helps!
I feel you—finding a good example dataset for panel data analysis can be a pain.

Try the "Panel Study of Income Dynamics" (PSID). It's got household data over decades, so perfect for practicing fixed/random effects models.

Also, Kaggle has some decent panel datasets if you search for "longitudinal data" or "time-series cross-section." Not all are perfect, but worth a look!
For a fun one, check out the "European Soccer Database" on Kaggle. It’s got player stats over multiple seasons—kinda cool if you’re into sports!

Otherwise, the "American Community Survey" (ACS) has panel-like data if you subset it right. Not perfectly clean, but real-world messy, which is good practice.

Pro tip: Stata’s built-in datasets (like "nlswork") are also solid for learning panel data analysis.
If you’re using R, the "plm" package comes with example dataset for panel data analysis, like "Grunfeld" (investment data) and "Produc" (state-level econ stuff). Super handy for testing code.

Otherwise, IPUMS has microdata with panel dimensions—kinda niche but super detailed.

P.S. Avoid the "Fatalities" dataset unless you wanna dive into heavy econometrics right away lol.
Yo, the "German Socio-Economic Panel" (SOEP) is *chef’s kiss* for this. It’s got everything—income, health, education—over years.

Also, if you’re into R, the "pder" package has neat datasets like "RiceFarms" for balanced panels.

Btw, avoid the "Baltagi" book datasets unless you’re ready for hardcore econometrics. They’re clean but boring af.
Wow, thanks for all the suggestions! The "National Longitudinal Survey of Youth" and "European Soccer Database" sound especially interesting—gonna try those first.

Quick follow-up: anyone know if the PSID requires special permissions to access? Saw it mentioned a few times but couldn’t tell if it’s fully open.

Also, big shoutout for the R package tips. Didn’t realize "plm" had built-in datasets—definitely checking that out too.

Y’all are lifesavers!
For something super relatable, try the "High School & Beyond" dataset. It tracks students over time—perfect for mixed-effects models.

Or, if you want macro, the "Maddison Project" has GDP data for countries over centuries (yes, centuries!).

Fair warning: some datasets need cleaning, but that’s part of the learning process, right?



Users browsing this thread: 1 Guest(s)