DistilKaggle
A more accessible corpus of computational notebooks.
What makes a large research dataset usable?
DistilKaggle combines large-scale data processing with computational notebook research. The project distills notebook content and code metrics into a corpus that other researchers can use to study how computational work is written and shared.
What connects the work
Infrastructure shapes the questions researchers can ask. This project treats making data usable as part of the research itself.
