Identifying bottlenecks in a Supply chain based on risk assessment for the whole chain. Having an inventory of possible choices selected to fit the design and expertise of the developers
A working prototype for a selected number of choices in the high risk areas of the project to test what works and to be prepared for any sudden breakdown of that link in the chain -> proposal
Selected 41 auth with 1K+ messages each
tagged by message content id
Result so far: Most similar messages typically not from the same author
Expanding to full set of authors
Combining messages of each author as paragraph tags
Method 2:
4M of redit but it crashes
Find a sample of users (with at least 1post with 200+ chars)
Find all comments associated with user and compare models among models
**Data Set 2: **
MongoDB: da1.eecs.utk.edu
Database.Collection: foreseer-reddit.comments
Objective 2: Investigate the possibility of using files modified to identify individuals
Investigate developer profile similarity function defined in, for
example:
Mockus, A. (2009, May). Succession: Measuring transfer of code and
developer productivity. In Proceedings of the 31st International
Conference on Software Engineering (pp. 67-77). IEEE Computer Society.
individual to project name on a subset two idx files
/home/yli118/identifying/java - reads delta.idx..gz -> top ten ids for any given id
plan: distance based on individual to file:
Objective 3: Summarize individuals in a scorecard
what statisics are collected: (duration, #changes, all commit time stamps)
what portion of data is processed: the entire delta.idx.*
how/where te results are stored:
/home/lwan1/{intervals,changes}.out
scorecard.py
still working on sentiment
Duration of activity (time from first to last change)
Number of changes, files, other people changing the same files,
number of projects
Skill: changes to files in different languages
Productivity growth
Uniformity of activity over time
Tone in text messages
LDA of text messages
Objective 4: Giant graph: get connected components
Experiment 1: get connected components using version history and
content id
Use Data Set 1
Also
Data Set 3:da3:/data/bkp/All.new.idx.*.gz
content id: 283056503
Size: 18147
File/version:
NewNewNew34.0/github.com_Velek_k-9.git/AndroidManifest.xml/4b9f21897ddb057fa94ee7b57d85985f2f2dad5f
Match to Data Set 1 using File/version
takes content id and project and removes content ids associated
with a single project
progress: of the 16 files 11 are done
output: /data/bkp/*.filtered
Once done: create connected components
Experiment 2: get connected components using version history, content
id, and authorship