By the same authors

Crossflow: A framework for distributed mining of software repositories

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Full text download(s)

Published copy (DOI)



Publication details

Title of host publicationProceedings - 2019 IEEE/ACM 16th International Conference on Mining Software Repositories, MSR 2019
DatePublished - 1 May 2019
Number of pages5
PublisherIEEE Computer Society Press
Original languageEnglish
ISBN (Electronic)9781728134123

Publication series

NameIEEE International Working Conference on Mining Software Repositories
ISSN (Print)2160-1852
ISSN (Electronic)2160-1860


Large-scale software repository mining typically requires substantial storage and computational resources, and often involves a large number of calls to (rate-limited) APIs such as those of GitHub and StackOverflow. This creates a growing need for distributed execution of repository mining programs to which remote collaborators can contribute computational and storage resources, as well as API quotas (ideally without sharing API access tokens or credentials). In this paper we introduce Crossflow, a novel framework for building distributed repository mining programs. We demonstrate how Crossflow can delegate mining jobs to remote workers and cache their results, and how workers can implement advanced behaviour such as load balancing and rejecting jobs they cannot perform (e.g. due to lack of space, credentials for a specific API).

    Research areas

  • Client-server systems, Computer aided software engineering, Data analysis, Data collection, Data flow computing, Data integration, Distributed processing, Modeling, Open source software, Pipeline pro cessing, Public domain software, Scalability, Software engineering

Discover related content

Find related publications, people, projects, datasets and more using interactive charts.

View graph of relations