Thursday, June 26, 2014
Alice Zheng's GraphLab- O'Reilly Webinar is now online
Sunday, June 22, 2014
Interesting taxi rides dataset
I got the following from my collaborator Zach Nation. NY taxi ride dataset that was not properly anonymized and was reverse engineered to find interesting insights in the data.
For the sport, I have used GraphLab Create to load and analyze this dataset. I started with an image of some NY taxis:
Using GraphLab Create I was able to reverse engineer the anonymizaiton and query the data based on the medallion number (for example 8J77 for the lower left taxi in the image).
I was further able to dig into personal details based on the medallion number:
And finally ask questions like how much money the taxis in the image made in a certain week?
Anyone who wants to try it out is welcome to email me, I can send you the ipython notebook to play with.
For the sport, I have used GraphLab Create to load and analyze this dataset. I started with an image of some NY taxis:
Using GraphLab Create I was able to reverse engineer the anonymizaiton and query the data based on the medallion number (for example 8J77 for the lower left taxi in the image).
I was further able to dig into personal details based on the medallion number:
And finally ask questions like how much money the taxis in the image made in a certain week?
Anyone who wants to try it out is welcome to email me, I can send you the ipython notebook to play with.
Monday, June 16, 2014
Be a detective with GraphLab create! Follow bitcoin money transactions to reveal a criminal!
Just got a note from my collaborator Brian Kent, who just related a new notebook which shows how to analyze Bitcoin money transactions using GraphLab Create. Using this notebook, Brian is trying to reveal a thief who stole 25,000$ Bitcoin money. Here is a graph of some of the thief transactions:
To learn the rest of the story you will need to read the full notebook.
Related blog posts: Graph analytics is a promising tools for fraud detection and security. Recently, Cisco announced that GraphLab is part of their security stack. PNNL is using GraphLab for its cyber security projects. Lab41 (US gov. research lab) combines Titan and GraphLab for a powerful social graph analytic tool.
To learn the rest of the story you will need to read the full notebook.
Related blog posts: Graph analytics is a promising tools for fraud detection and security. Recently, Cisco announced that GraphLab is part of their security stack. PNNL is using GraphLab for its cyber security projects. Lab41 (US gov. research lab) combines Titan and GraphLab for a powerful social graph analytic tool.
Saturday, June 14, 2014
Community detection survey by Lab41
Just got my hands on the community detection survey made by Lab41. A very comprehensive overview of the popular and useful methods to know. Some of the included methods are Girwan Newman, Infomaps, Fast Unfolding, Cesna and many more.
One of the interesting algorithms is BigClam:
One of the interesting algorithms is BigClam:
Friday, June 13, 2014
Lab41 releases open source code for GraphLab + Yarn integration
Just heard from Erik Tryzlaar from Lab41, that a new github open source project called Twill is alive. The project allows for running GraphLab tasks on a Hadoop 2.0 cluster which supports Yarn.
To remind, Pivotal have also their own wrapper which allows for running GraphLab on their Hadoop cluster, as part of their HD project.
A lot of exciting activities from different parties who are helping to make Graphlab Hadoop compatible! We will also release some news from GraphLab about this direction soon.
To remind, Pivotal have also their own wrapper which allows for running GraphLab on their Hadoop cluster, as part of their HD project.
A lot of exciting activities from different parties who are helping to make Graphlab Hadoop compatible! We will also release some news from GraphLab about this direction soon.
Fantastic talk by Dafna Shahaf - Stanford
This week I attended a great talk by Dafna Shahaf. In a nutshell, she has a method for finding surprising insights in the data. An open source project is on the way for sharing some of those tools.
Several applications domains were covered. For example, in the medical domain, two of the system findings (out of 4) are major medical breakthrough revelation as defined by external physician who examined the output. For commerce, the system can find surprising Amazon products to recommend to people. Here is a nice example:
For people who are looking for child's toys there are a lot of related selections in the pet section. For people who need a bath mat, there are related products in the car department which are much cheaper..
Several applications domains were covered. For example, in the medical domain, two of the system findings (out of 4) are major medical breakthrough revelation as defined by external physician who examined the output. For commerce, the system can find surprising Amazon products to recommend to people. Here is a nice example:
For people who are looking for child's toys there are a lot of related selections in the pet section. For people who need a bath mat, there are related products in the car department which are much cheaper..
Leading Cancer Hospital utilizes GraphLab LDA for HealthCare
Just heard very interesting report from Xinghua Lou, a researcher of machine learning in Microsoft Research. Xinghua utilized GraphLab topic modeling for clustering health related documents. This work was reported at the big data innovation summit 2014.
From the KDnuggets blog post about this work:
"Among various techniques for understanding text corpus, we chose LDA topic models (implemented in GraphLab) because of its previous success in understanding scientific literature as well as webpages. We followed a process roughly as follows: data cleaning and standardization, topic modeling, clinical note clustering and visualization, community finding and cancer-gene correlation analysis. This process was mainly implemented by Katherine Chanunder my supervision. We had a few interesting findings, such as a community of patients who highly care about the risk of the treatment, the ability of predicting icd-9 code from topic modeling output, and some interesting correlations between patient profile and genetic mutation tests (some supported by previous published research)."
From the KDnuggets blog post about this work:
"Among various techniques for understanding text corpus, we chose LDA topic models (implemented in GraphLab) because of its previous success in understanding scientific literature as well as webpages. We followed a process roughly as follows: data cleaning and standardization, topic modeling, clinical note clustering and visualization, community finding and cancer-gene correlation analysis. This process was mainly implemented by Katherine Chanunder my supervision. We had a few interesting findings, such as a community of patients who highly care about the risk of the treatment, the ability of predicting icd-9 code from topic modeling output, and some interesting correlations between patient profile and genetic mutation tests (some supported by previous published research)."
Subscribe to:
Posts (Atom)




