Since 2004, FLOSSmole aims to:
provide a community for researchers to discuss public data about FLOSS development. FLOSSmole contains: Nearly 1 TB of data covering the period 2004-now, and growing with data sets from over 275 web spidering operations, and growing each month. This includes data about more than 500,000 different open source projects and their developers. How to Cite FLOSSmole Data All original data is copyright of its owners.
Please be aware that your use of this data for research may require approval by your company's or institution's IRB (Institutional Review Board). |
|||
We're cruising ahead with January 2012 releases. Grab the data from Google Code site or from the teragrid. Freecode - done (formerly known as Freshmeat) Google Code - still running Free Software Foundation - bug still not fixed (this is my fault) #51 Interesting things: most popular data from November ..... drumroll please.... Google Code, Github. |
|||
Google Code is our longest data collection effort each month. We've collected everything for November and posted it for your data mining pleasure. Get the files or access it on the Teragrid with direct database access (datasource_id=285). |
|||
Three things happened recently to affect our Freshmeat collection 1. Freshmeat announced a name change to Freecode. What I've done is as follows: For issue #1 - decided not to rename our abbreviation for Freshmeat. It will remain "FM". For issue #2 & 3 - Added a new table to hold the tags associated with a project. It's called fm_projects_tags. CREATE TABLE IF NOT EXISTS `fm_project_tags` ( Added a new release file to hold the data from this table. The new file is called fmProjectTags2011-Nov.txt. Did not remove trove; we are still collecting the trove. Although there is no longer any "trove definition" list that I know of to describe each trove number, so these are not as useful as the "tags". But I'm leaving this alone in the database for historical purposes. Here is a shot of the tags page for a sample project on Freshmeat, called amms.
Here is a shot of the way the tags look now in our release files (or database table) for that same project (#78922)
|
|||
Here is the status of the November 2011 collection: done & ready to download on Google Code or query in Teragrid... still collecting... collectors broken and waiting to be fixed... |
|||
One of the papers at the 2011 OSS conference is entitled "Building Knowledge in Open Source Software Research in Six Years of Conferences". It surveys the contributions of papers presented at the OSS conferences, and builds social networks of the papers, identifying research streams along the way. Findings particular to FLOSSmole: "Cluster #82. The largest cluster originates from node #82. Paper #82 introduces the OSSmole project (later called FLOSSmole). OSSmole is a repository of data, scripts, and analysis of data collected from OSS projects." and "Large clusters are initiated by empirical papers with the only exception being the paper on the FLOSSmole repository." and "Papers with a large number of citations [ed: such as FLOSSmole paper] are synthesizers of research often presenting a framework or a platform to guide research in OSS." and "In particular, we have found that the creation of a big repository for data mining (FLOSSmole) has originated research in social network analysis, tools for data mining, and analysis of code artefacts to understand maintenance processes, specifically bug fixing." |
|||
1. Free Software Foundation directory changed their layout to a wiki so we're re-writing our collector to parse RDF instead. This will change the tables we use for FSF data now. 2. We were able to convince our dear colleague Audris Mockus to run his Google Code collector and gather the latest list of project names for us. SWEET! This means a Google Code run is imminent. 3. UDD and Debian still need to be re-run, and automated. 4. In case you are keeping track of the different forges, Berlios is shutting down as of Dec 31 2011. We're still plugging along with all this stuff. Hope you are finding the data helpful. Let us know what we can provide. (and join the Mailing List!) |
|||
Here is the pre-print copy of the paper on forges that David and I have written. I am going to present at HICSS 45 in January. Squire, M. and Williams, D. (2012). Describing the software forge ecosystem. 45th Hawaii International Conference on System Sciences. Maui, Hawaii. January 4-7. Forthcoming. |
|||
Here is the status of each collection for September 2011: The stages are UPDATED as of 05-Sep-2011 at 12:41PM: Rubyforge - files released to Google Code & data uploaded into Teragrid Objectweb - files released to Google Code & data uploaded into Teragrid Free Software Foundation Directory - files released to Google Code & data uploaded into Teragrid Savannah - files released to Google Code & data uploaded into Teragrid Github - files released to Google Code & data uploaded into Teragrid Tigris - files released to Google Code & data uploaded into Teragrid Google Code - I am looking at getting a new list of projects as the one we've been using is quite old now (Oct 2010) Launchpad - files released to Google & data uploaded to Teragrid Alioth - files released to Google Code & data uploaded to Teragrid (new tables made) Debian Metrics - waiting on README Ultimate Debian Database - importing into database; error on table create |
|||
Summer is a beautiful thing. Moles, we've got a huge Google Code release for you (ds=271), and the re-vamped Launchpad (ds=272), and also Github (ds=273). Get your FRESH June data on our Google Code Downloads Page or LIVE on the Teragrid. Tigris is fixed and is running right now. We're also writing a new collector for Alioth! Lots of new stuff. Got a bug in the Freshmeat collector, so I'm wrangling that. Thanks to a user for reporting that bug. Don't forget we do have a bug-tracking system on Google Code. Finally, we've got a fresh UDD upload and Debian data coming soon also. We're just so productive right now! Also don't forget to check out our collection of Everything You Ever Wanted to Know About Code Forges - data also available on our Google Code download site. |
|||


