Collection information

Details about the repository collections

September 2012 data released

Data files have been released for September 2012. Go check it out on our Google Code downloads page or sign up for direct database access.

Special release notes:
--The Google Code, Github, and Launchpad collections are not included this month.
--Alioth is back, and Tigris is back, including emails.

Data Resources: 

August 2012 data released

Data files have been released for August 2012. Go check it out on our Google Code downloads page or sign up for direct database access.

Special release notes:
--The Google Code, Github, and Launchpad collections are not included this month.

Data Resources: 

July 2012 data released

Data files have been released for July 2012. Go check it out on our Google Code downloads page or sign up for direct database access.

Data Resources: 

May 2012 data releases

Data files have been released for May 2012. Go check it out on our Google Code downloads page.

Re-writes:
--Free Software Foundation has been re-written from scratch to match their new layout.
--Google Code collector has been re-written to fix a few bugs (still running)
--Launchpad has been finished and will be re-written for June to fix bugs
--Alioth is being re-written to fix bugs

Data Resources: 

February Github data released

February data has been released for Github.

Get the data here from our Google Code downloads page or request direct database access here.

Included with Github data are the following values:
project name
developer name
description
private yes/no
fork number
homepage
number of watchers
open issues
...and all the xml values that these fields are based on!

Have fun!

Data Resources: 
Tags: 

February Google Code data released

Google Code data has been released for January/February 2012.

Get the data here from our Google Code downloads page or request direct database access here.

Be aware that there is one open bug for Google Code collection that may affect your use of this data.

Data Resources: 
Tags: 

January 2012 releases

We're cruising ahead with January 2012 releases. Grab the data from Google Code site or from the teragrid.

Freecode - done (formerly known as Freshmeat)
Savannah - done
Tigris - done
Rubyforge - done
Objectweb - done
Launchpad - done

Data Resources: 

Google Code data available

Google Code is our longest data collection effort each month. We've collected everything for November and posted it for your data mining pleasure. Get the files or access it on the Teragrid with direct database access (datasource_id=285).

Data Resources: 

November 2011 data entered

Here is the status of the November 2011 collection:

done & ready to download on Google Code or query in Teragrid...
============
RUBYFORGE
OBJECTWEB
TIGRIS
LAUNCHPAD
SAVANNAH
ALIOTH
GITHUB

still collecting...
============
GOOGLE

Data Resources: 

Current challenges for Fall

1. Free Software Foundation directory changed their layout to a wiki so we're re-writing our collector to parse RDF instead. This will change the tables we use for FSF data now.

2. We were able to convince our dear colleague Audris Mockus to run his Google Code collector and gather the latest list of project names for us. SWEET! This means a Google Code run is imminent.

3. UDD and Debian still need to be re-run, and automated.

Data Resources: 

Pages