<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Datablend &#187; gephi</title>
	<atom:link href="http://datablend.be/?cat=30&#038;feed=rss2" rel="self" type="application/rss+xml" />
	<link>http://datablend.be</link>
	<description>Big Data Simplified</description>
	<lastBuildDate>Mon, 07 Sep 2015 09:04:17 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>hourly</sy:updatePeriod>
	<sy:updateFrequency>1</sy:updateFrequency>
	<generator>http://wordpress.org/?v=3.6.1</generator>
		<item>
		<title>Coalition-Cocktail &#8211; Hacking the Elections @ Engagor</title>
		<link>http://datablend.be/?p=440</link>
		<comments>http://datablend.be/?p=440#comments</comments>
		<pubDate>Tue, 27 May 2014 14:57:42 +0000</pubDate>
		<dc:creator>Davy Suvee</dc:creator>
				<category><![CDATA[circos]]></category>
		<category><![CDATA[engagor]]></category>
		<category><![CDATA[gephi]]></category>
		<category><![CDATA[graph]]></category>
		<category><![CDATA[hackaton]]></category>
		<category><![CDATA[vk14]]></category>

		<guid isPermaLink="false">http://datablend.be/?p=440</guid>
		<description><![CDATA[Last weekend, Engagor organised their hacktheelections hackaton. The Datablend team (Quentin, Stijn and Davy) was joined by Marc Broos, Tim Coene and Josbert van de Zande with one goal in mind: trying to visualise the (pre-arranged?) political coalition and, if possible, also predict the formation-period. Technically, we extracted over 160K tweets through the Engagor API.<p><a href="http://datablend.be/?p=440">Continue Reading →</a></p>]]></description>
				<content:encoded><![CDATA[<p style="text-align: left;">Last weekend, <a title="Engagor" href="https://engagor.com" target="_blank">Engagor</a> organised their <a title="hacktheelections" href="http://www.hacktheelections.com" target="_blank">hacktheelections</a> hackaton. The Datablend team (Quentin, Stijn and Davy) was joined by Marc Broos, Tim Coene and Josbert van de Zande with one goal in mind: trying to visualise the (pre-arranged?) political coalition and, if possible, also predict the formation-period.</p>
<p style="text-align: left;">Technically, we extracted over 160K tweets through the Engagor API. Next, a &#8220;sentiment&#8221;-based political graph was build and stored in the <a href="http://www.neo4j.org" title="Neo4J" target="_blank">Neo4J</a> graph database. A <a href="http://www.gephi.org" title="Gephi" target="_blank">Gephi</a>-based visualisation, based upon community detection, revealed the &#8220;truth&#8221;, which was, to be honest, a bit disappointing and at the same time somewhat expected: the political parties from Wallionia form a solid community while the Flemish parties are interconnected through many clusters. A warning sign for the upcoming formation process?</p>
<p style="text-align: left;">The slidedeck below provides an overview of the single day of hacking. Our results ranked third (on 8 teams). Although being very informative, it&#8217;s quite hard to compete with a &#8220;pokemon&#8221;-themed fighting animation between politicians *wink*. Many thanks to <a title="Engagor" href="https://engagor.com" target="_blank">Engagor</a> for the spotless organisation!</p>
<p><center><iframe src="http://www.slideshare.net/slideshow/embed_code/35157124" width="600" height="489" frameborder="0" marginwidth="0" marginheight="0" scrolling="no"></iframe><br/><br/></center></p>
<p></p>]]></content:encoded>
			<wfw:commentRss>http://datablend.be/?feed=rss2&#038;p=440</wfw:commentRss>
		<slash:comments>0</slash:comments>
		</item>
		<item>
		<title>The power of graphs to analyse biological data</title>
		<link>http://datablend.be/?p=344</link>
		<comments>http://datablend.be/?p=344#comments</comments>
		<pubDate>Mon, 02 Dec 2013 07:04:34 +0000</pubDate>
		<dc:creator>Davy Suvee</dc:creator>
				<category><![CDATA[gephi]]></category>
		<category><![CDATA[graph]]></category>
		<category><![CDATA[graphconnect]]></category>
		<category><![CDATA[mongodb]]></category>
		<category><![CDATA[neo4j]]></category>
		<category><![CDATA[NoSQL]]></category>
		<category><![CDATA[visualisation]]></category>

		<guid isPermaLink="false">http://datablend.be/?p=344</guid>
		<description><![CDATA[Watch Davy Suvee present at GraphConnect London 2013 on the power of graph databases to analyse biological datasets. The Power of Graphs to Analyze Biological Data &#8211; Davy Suvee @ GraphConnect London 2013 from Neo Technology on Vimeo.]]></description>
				<content:encoded><![CDATA[<p>Watch Davy Suvee present at <a href="http://www.graphconnect.com/london/" title="GraphConnect London 2013">GraphConnect London 2013</a> on the power of graph databases to analyse biological datasets.</p>
<p><iframe src="//player.vimeo.com/video/80463932" width="500" height="281" frameborder="0" webkitallowfullscreen mozallowfullscreen allowfullscreen></iframe>
<p><a href="http://vimeo.com/80463932">The Power of Graphs to Analyze Biological Data &#8211; Davy Suvee @ GraphConnect London 2013</a> from <a href="http://vimeo.com/neo4j">Neo Technology</a> on <a href="https://vimeo.com">Vimeo</a>.</p>
<p></p>]]></content:encoded>
			<wfw:commentRss>http://datablend.be/?feed=rss2&#038;p=344</wfw:commentRss>
		<slash:comments>0</slash:comments>
		</item>
		<item>
		<title>Yelp graph: checkin-based business clustering</title>
		<link>http://datablend.be/?p=308</link>
		<comments>http://datablend.be/?p=308#comments</comments>
		<pubDate>Sun, 01 Dec 2013 11:51:58 +0000</pubDate>
		<dc:creator>Davy Suvee</dc:creator>
				<category><![CDATA[gephi]]></category>
		<category><![CDATA[graph]]></category>
		<category><![CDATA[neo4j]]></category>
		<category><![CDATA[visualisation]]></category>
		<category><![CDATA[yelp]]></category>

		<guid isPermaLink="false">http://datablend.be/?p=308</guid>
		<description><![CDATA[Recently, Yelp made available a sample dataset from the greater Phoenix metropolitan area including around 11.000 business, 8000 checkin-sets, 43.000 users and 230.000 user reviews. With the help of this data, data scientists can execute real-life experiments with various data mining/machine learning algorithms. In our case, we are interested in finding out whether it is possible<p><a href="http://datablend.be/?p=308">Continue Reading →</a></p>]]></description>
				<content:encoded><![CDATA[<p style="text-align: justify;">Recently, <a title="yelp" href="http://yelp.com" target="_blank">Yelp</a> made available <a title="a sample dataset" href="http://www.yelp.co.uk/dataset_challenge" target="_blank">a sample dataset</a> from the greater Phoenix metropolitan area including around 11.000 business, 8000 checkin-sets, 43.000 users and 230.000 user reviews. With the help of this data, data scientists can execute real-life experiments with various data mining/machine learning algorithms. In our case, we are interested in finding out whether it is possible to visually cluster businesses by category, based purely on their checkin data. The checkin data itself is available on a day-hour level: for each business, it is possible to retrieve the number of checkins on a Sunday afternoon between 3 and 4. So, with only this data in mind, are we able to cluster businesses as being restaurants or fashion stores, based purely on the correlations calculated amongst their checkin data? For this experiment, we use the <a title="Neo4J" href="http://neo4j.org" target="_blank">Neo4J</a> graph database for storing our checkin-based correlation graph and employ the <a title="Gephi" href="http://gephi.org" target="_blank">Gephi</a> graph visualisation platform for interpreting the identified business communities/clusters. As always, the <a title="full source code" href="https://github.com/datablend/yelp-graph" target="_blank">full source code</a> of this article can be found on the Datablend public github repository (although you will need to acquire the dataset yourself through the <a title="Yelp Dataset Challenge" href="http://www.yelp.co.uk/dataset_challenge" target="_blank">Yelp Dataset Challenge</a> portal).</p>
<h3>1. Building the Neo4J checkin correlation graph</h3>
<p style="text-align: justify;">We start by parsing both the business and checkin json-files from the Yelp Dataset challenge. Unfortunately, checkin data is available for only 8,282 out of the 11,537 supplied businesses. In addition, many of these have only a limited set of associated checkins. Hence, in order to make sure that only relevant correlations are calculated, we ignore the ones that have less than a 100 checkins, resulting in around 1920 remaining businesses.</p>
<p style="text-align: justify;">Next, we try to identify the correlation between two businesses by using the Pearson Correlation Coefficient (read <a title="this site" href="https://statistics.laerd.com/statistical-guides/pearson-correlation-coefficient-statistical-guide.php" target="_blank">this site</a> for a nice introduction). Simply put, we try to identify whether a linear association exists between the checkins of two individual businesses. In our case, the calculation is based upon 168 data points (24 hours x 7 days), the idea being that two breakfast restaurants will get most of their checkins from the morning till noon, while two bars will get most of their checkins during the evening and at night. Hence, we expect the correlation between two businesses of the same type to be quite high, while different types of businesses (i.e. a breakfast restaurant and a bar) will  result in little or no correlation.</p>
<p style="text-align: justify;">Time to get our hands dirty. After parsing the data files, we use the existing apache.commons.math to calculate the pairwise correlation between the checkin datasets of the 1920 businesses. If the resulting coefficient is 0.8 or higher, we consider both businesses to be correlated. We create a unique node for each business within the Neo4J graph and combine them via a &#8220;correlated&#8221;-relationship.</p>
<script src="https://gist.github.com/7704584.js"></script>
<p style="text-align: justify;">The generated graph contains 606 unique nodes (i.e. businesses that are correlated to at least one other business) and 2585 edges (i.e. actual correlations).</p>
<h3>2. Gephi interpretation</h3>
<p style="text-align: justify;">Our next task is to observe whether groups of businesses exists that are highly correlated (i.e. highly interconnected) and identify whether these correlations makes sense. In order to do so, we import our Neo4J correlation graph in Gephi through the <a title="Gephi Neo4J plugin" href="https://marketplace.gephi.org/plugin/neo4j-graph-database-support/" target="_blank">Gephi Neo4J plugin</a>. Once loaded, we run the <a title="modularity" href="http://wiki.gephi.org/index.php/Modularity" target="_blank">modularity</a>-function to identify meaningful communities. These computed communities are then used to partition (i.e. color) the nodes (and their related edges) so that clusters can easily be observed. Next, we apply K-core filtering, in our case 3-core, to keep the subgraph from which all nodes have a degree of at least 3 (i.e. 3 relationships with other nodes). The size of the nodes (and their associated labels) is configured to be proportional with their degree. Finally, we apply <a title="Fruchterman-Reingold" href="http://wiki.gephi.org/index.php/Fruchterman-Reingold" target="_blank">Fruchterman-Reingold</a> lay-outing in order to clearly visualise the various clusters.</p>
<p style="text-align: justify;"><a href="http://datablend.be/wp-content/uploads/2013/12/yelp-graph.jpg" target="_blank"><img class="alignnone size-medium wp-image-334" alt="yelp-graph" src="http://datablend.be/wp-content/uploads/2013/12/yelp-graph.jpg"/></a></p>
<p style="text-align: justify;">We can easily observe 8 communities, but are these clusters meaningful? The pink cluster on the right-end side is highly interconnected (i.e. all nodes of the cluster have mutual correlations). Most of them can be identified as being breakfast diners (ex. <a href="http://www.yelp.co.uk/search?find_desc=the+good+egg&#038;find_loc=Phoenix%2C+AZ%2C+USA" title="The Good Egg">The good egg</a>, <a href="http://www.yelp.co.uk/biz/the-breakfast-joynt-scottsdale-2" title="The breakfast joynt">The breakfast joynt</a> and <a href="http://www.yelp.co.uk/biz/orange-table-scottsdale" title="Orange table">Orange table</a>). Cool. This certainly make sense, as most of these business have checkins early morning until early afternoon. The yellow cluster on the top contains various department stores (including <a href="http://www.yelp.co.uk/biz/costco-phoenix-4" title="Costco">Costco</a>, <a href="http://www.yelp.co.uk/biz/nordstrom-rack-phoenix-2#query:nordstroms%20rack" title="Nordstrom Rack" target="_blank">Nordstrom Rack</a> and <a href="http://www.yelp.co.uk/biz/ikea-tempe" title="IKEA" target="_blank">IKEA</a>). Again meaningful, as most of them open their doors somewhere around 10AM and close around 7PM. At first sight, it seems strange that the coffee places are correlated into two separate groups (yellow cluster at the bottom and pink cluster on the top). The reason however is simple: some of them close late afternoon while others are open until midnight.</p>
<h3>3. Conclusion</h3>
<p style="text-align: justify;">The Neo4J/Gephi solution works remarkably well to visually identify the various business clusters from the Yelp dataset. In a next blog article, we will show how to use the <a href="http://en.wikipedia.org/wiki/K-nearest_neighbors_algorithm" title="k-nearest neighbours algorithm">k-nearest neighbours algorithm</a> to automatically predict the type of business based upon solely the checkin information.</p>
<p></p>]]></content:encoded>
			<wfw:commentRss>http://datablend.be/?feed=rss2&#038;p=308</wfw:commentRss>
		<slash:comments>0</slash:comments>
		</item>
		<item>
		<title>Running along the graph using Neo4J Spatial and Gephi</title>
		<link>http://datablend.be/?p=262</link>
		<comments>http://datablend.be/?p=262#comments</comments>
		<pubDate>Wed, 04 Jan 2012 09:48:13 +0000</pubDate>
		<dc:creator>Davy Suvee</dc:creator>
				<category><![CDATA[gephi]]></category>
		<category><![CDATA[neo4j]]></category>
		<category><![CDATA[spatial]]></category>
		<category><![CDATA[visualisation]]></category>

		<guid isPermaLink="false">http://datablend22.lin3.nucleus.be/?p=262</guid>
		<description><![CDATA[When I started running some years ago, I bought a Garmin Forerunner 405. It&#8217;s a nifty little device that tracks GPS coordinates while you are running. After a run, the device can be synchronized by uploading your data to the Garmin Connect website. Based upon the tracked time and GPS coordinates, the Garmin Connect website<p><a href="http://datablend.be/?p=262">Continue Reading →</a></p>]]></description>
				<content:encoded><![CDATA[<p style="text-align: justify;">When I started running some years ago, I bought a <a target='_blank' href="https://buy.garmin.com/shop/shop.do?pID=11039&#038;ra=true#owners">Garmin Forerunner 405</a>. It&#8217;s a nifty little device that tracks GPS coordinates while you are running. After a run, the device can be synchronized by uploading your data to the <a target='_blank' href="http://connect.garmin.com">Garmin Connect website</a>. Based upon the tracked time and GPS coordinates, the Garmin Connect website provides you with a detailed overview of your run, including <span class="highlight"><em>distance</em></span>, <span class="highlight"><em>average pace</em></span>, <span class="highlight"><em>elevation loss/gain</em></span> and <span class="highlight"><em>lap splits</em></span>. It also visualizes your run, by overlaying the tracked course on Bing and/or Google maps. Pretty cool! One of my last runs can be found <a target='_blank' href="http://connect.garmin.com/activity/138373187">here</a>.</p>
<p style="text-align: justify;">Apart from <span class="highlight"><em>simple aggregations</em></span> such as total distance and average speed, the Garmin Connect website provides little or no support to gain deeper insights in all of my runs. As I often run the same course, it would be interesting to calculate my <span class="highlight"><em>average pace at specific locations</em></span>. When combining the data of all of my courses, I could deduct <span class="highlight"><em>frequently encountered locations</em></span>. Finally, could there be a <span class="highlight"><em> correlation</em></span> between my <span class="highlight"><em>average pace</em></span> and my <span class="highlight"><em>distance from home?</em></span> In order to come up with answers to these questions, I will import my running data into a <a target='_blank' href="https://github.com/neo4j/spatial">Neo4J Spatial</a> datastore. Neo4J Spatial extends the <a target='_blank' href="http://neo4j.org/">Neo4J Graph Database</a> with the necessary tools and utilities to store and query spatial data in your graph models. For visualizing my running data, I will make use of <a target='_blank' href="http://gephi.org/">Gephi</a>, an open-source visualization and manipulation tool that allows users to interactively browse and explore graphs.</p>
<p>&nbsp;</p>
<h3>1. Extracting GPX data</h3>
<p style="text-align: justify;">The Garmin Connect website allows to download running data through various formats, including <span class="highlight"><em>KML</em></span>, <span class="highlight"><em>TCX</em></span> and <span class="highlight"><em>GPX</em></span>. <a target='_blank' href="http://topografix.com/gpx.asp">GPX</a> (the GPS Exchange Format) is a light-weight XML data format that is used for interchanging GPS data (waypoints, routes, and tracks) between applications and web services. Below, you can find a GPX extract enumerating several tracked points. Each of these points contains the <span class="highlight"><em>GPS location</em></span>, the <span class="highlight"><em>elevation</em></span> and the corresponding <span class="highlight"><em>timestamp</em></span>.</p>
<script src="https://gist.github.com/1559458.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">Based upon this data, one is able to calculate various metrics, including <span class="highlight"><em>pace</em></span>. For this, we will use <a target='_blank' href="http://gpstools.sourceforge.net/">GPSdings</a>, a Java library that provides the required functionality to extract and analyze GPX data. We start by reading in a GPX file. Afterwards, we <span class="highlight"><em>analyze</em></span> the content using the GPSdings <em>TrackAnalyzer</em> which, amongst other metrics, calculates the pace for each point that was tracked during a run. The information we need is stored in the first segment of the first track.</p>
<script src="https://gist.github.com/1559808.js"></script>
<p>&nbsp;</p>
<h3>2. Importing GPS data in Neo4J Spatial</h3>
<p style="text-align: justify;"><span class="highlight"><em>Neo4J Spatial</em></span> is build on top of <span class="highlight"><em>Neo4J</em></span> and provides support for <span class="highlight"><em>spatial data</em></span>. Once your data is stored, <span class="highlight"><em>spatial operations</em></span> can be executed, which for instance allow to search for data within specified regions or within a specified distance of a particular point of interest. We start by setting up a Neo4J <em>EmbeddedGraphDatabase</em>. We then wrap it as a <em>SpatialDatabaseService</em>, which allows us to create an <em>EditableLayer</em>. <em>EditableLayer</em> is Neo4J&#8217;s main abstraction, which is used to define a <span class="highlight"><em>collection of geometries</em></span>. Each layer needs to be initialized with a specific <em>GeometryEncoder</em>, which acts a kind of adapter to map from the graph to the geometries and vice versa. In our case, we will employ the <em>SimplePointEncoder</em>.</p>
<script src="https://gist.github.com/1559893.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">Adding spatial data to the running layer is very easy. We start by creating a <em>Coordinate</em> for each point that is parsed by GPSdings. Next, we add this new coordinate to the running layer. This operation returns a <em>SpatialDatabaseRecord</em> which, under the hood, is just a <span class="highlight"><em>regular Neo4J node</em></span>. Hence, we can add any property we want to this node. In our case, we will add two properties. One property, named <span class="highlight"><em>speed</em></span>, indicating the (average) pace. One property, named <span class="highlight"><em>occurrences</em></span>, indicating the number of times this particular coordinate was encountered in the overall data set. Once the new coordinate is created, we connect the previous node with the newly created node through the <span class="highlight"><em>NEXT</em></span> relationship type. Hence, our graph is an <span class="highlight"><em>enumeration</em></span> of the encountered coordinates, <span class="highlight"><em>interlinked</em></span> through NEXT edges.</p>
<script src="https://gist.github.com/1559954.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">In case a coordinate is encountered multiple times, we <span class="highlight"><em>recalculate the average speed</em></span> and <span class="highlight"><em>increment the number of encounters</em></span>.</p>
<script src="https://gist.github.com/1560142.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">Unfortunately, chances are low to encounter an already existing coordinate, as coordinates in a GPX file have a 15-digit precision right of the decimal point. Instead of trying to <span class="highlight"><em>round</em></span> these coordinates ourselves, we will use the <span class="highlight"><em>Neo4J Spatial querying API</em></span>. A simple <span class="highlight"><em>nearest neighbor</em></span>-search limited to 20 meters allows us to find matching coordinates. (I choose 20 meters, as 20 is a little above the average distance between two coordinates). In case we find a coordinate within this 20-meter range, we will <span class="highlight"><em>reuse</em></span> it. Otherwise, we just create a <span class="highlight"><em>new coordinate</em></span>. The full algorithm for importing multiple GPX datasets can be found below.</p>
<script src="https://gist.github.com/1560218.js"></script>
<p>&nbsp;</p>
<h3>3. Visualizing running data</h3>
<p style="text-align: justify;">By using the <span class="highlight"><em>Neo4J Spatial querying API</em></span>, we are able to retrieve the set of coordinates that satisfy a particular condition. However, coordinates are somewhat <span class="highlight"><em>abstract</em></span> to interpret. Instead, we will use the excellent Gephi Graph visualization and exploration tool. By installing the <a target='_blank' href="http://gephi.org/tag/neo4j/">Gephi Neo4J plugin</a>, we are able to load and explore graphs that are stored in a Neo4J (Spatial) datastore. Let&#8217;s start by <span class="highlight"><em>importing</em></span>  our dataset in Gephi.</p>
<p><a target='_blank' href="http://datablend.be/wp-content/uploads/geo1.jpg">
<p align="center"><img width="550" src="http://datablend.be/wp-content/uploads/geo1.jpg" alt="gephi" /></p>
<p></a></p>
<p style="text-align: justify;">The displayed graph contains other types of nodes and edges (i.e. <em>Layer </em>and <em>RTree </em>index information), in addition to the coordinates and NEXT edges that we added ourselves. Let&#8217;s get rid of those by <span class="highlight"><em>filtering our graph</em></span> on the NEXT relationship-type.</p>
<p><a target='_blank' href="http://datablend.be/wp-content/uploads/geo2.jpg">
<p align="center"><img width="550" src="http://datablend.be/wp-content/uploads/geo2.jpg" alt="gephi" /></p>
<p></a>  </p>
<p style="text-align: justify;">Only half of the edges remain &#8230; However, we will still not gain novel insights from this mess. Let&#8217;s layout our graph by using the <a target='_blank' href="http://gephi.org/plugins/geolayout/">Gephi GeoLayout plugin</a>. This layouter takes <span class="highlight"><em>geocoded graphs</em></span> as input and will layout graphs according to the geocoded attributes. Make sure to increase scaling, as our coordinates are located closely together. Cool! This view clearly outlines the courses I&#8217;m running.</p>
<p><a target='_blank' href="http://datablend.be/wp-content/uploads/geo3.jpg">
<p align="center"><img width="550" src="http://datablend.be/wp-content/uploads/geo3.jpg" alt="gephi" /></p>
<p></a></p>
<p style="text-align: justify;">Let&#8217;s visualize the coordinates that were <span class="highlight"><em>frequently encountered</em></span> during the 4 runs that are imported in the Neo4J Spatial datastore. For this, we will use the <span class="highlight"><em>InDegree</em></span> node property, which indicates <span class="highlight"><em>the number of incoming edges</em></span> for each coordinate. We rank <span class="highlight"><em>node weight</em></span> (i.e. node size) through this property. Hence, frequently encountered nodes will show up bigger. In my case, frequently encountered coordinates are found around the place where I live (and hence start my runs) and on street intersections.</p>
<p><a target='_blank' href="http://datablend.be/wp-content/uploads/geo4.jpg">
<p align="center"><img width="550" src="http://datablend.be/wp-content/uploads/geo4.jpg" alt="gephi" /></p>
<p></a></p>
<p style="text-align: justify;">Let&#8217;s do one final analysis, namely a visualization that illustrates the <span class="highlight"><em>average pace throughout all runs</em></span>. For this, we rank both <span class="highlight"><em>node weight</em></span> and <span class="highlight"><em>node color</em></span> through the <span class="highlight"><em>speed</em></span> property. Hence, coordinates with a high average pace are colored green and show up bigger. Coordinates with a low average pace are colored red and show up smaller. With the blink of an eye, I can now interpret my average pace, taking into account my overall running data set!</p>
<p><a target='_blank' href="http://datablend.be/wp-content/uploads/geo5.jpg">
<p align="center"><img width="550" src="http://datablend.be/wp-content/uploads/geo5.jpg" alt="gephi" /></p>
<p></a></p>
<p>&nbsp;</p>
<h3>4. Conclusion</h3>
<p style="text-align: justify;">This article describes the use of the <span class="highlight"><em>Neo4J Spatial datastore</em></span> and <span class="highlight"><em>Gephi</em></span> to analyze Garmin running data. As always, the complete source code can be found on the <a target='_blank' href="https://github.com/datablend/neo4j-spatial-running">Datablend public GitHub repository</a>. Any ideas for other types of analysis that could be performed on the dataset?</p>
<p></p>]]></content:encoded>
			<wfw:commentRss>http://datablend.be/?feed=rss2&#038;p=262</wfw:commentRss>
		<slash:comments>0</slash:comments>
		</item>
		<item>
		<title>Visualizing RDF Schema inferencing through Neo4J, Tinkerpop, Sail and Gephi</title>
		<link>http://datablend.be/?p=260</link>
		<comments>http://datablend.be/?p=260#comments</comments>
		<pubDate>Mon, 21 Nov 2011 09:47:11 +0000</pubDate>
		<dc:creator>Davy Suvee</dc:creator>
				<category><![CDATA[gephi]]></category>
		<category><![CDATA[neo4j]]></category>
		<category><![CDATA[NoSQL]]></category>
		<category><![CDATA[rdf]]></category>
		<category><![CDATA[sail]]></category>

		<guid isPermaLink="false">http://datablend22.lin3.nucleus.be/?p=260</guid>
		<description><![CDATA[Last week, the Neo4J plugin for Gephi was released. Gephi is an open-source visualization and manipulation tool that allows users to interactively browse and explore graphs. The graphs themselves can be loaded through a variety of file formats. Thanks to Martin Škurla, it is now possible to load and lazily explore graphs that are stored<p><a href="http://datablend.be/?p=260">Continue Reading →</a></p>]]></description>
				<content:encoded><![CDATA[<p style="text-align: justify;">Last week, the <a target='_blank' href="http://neo4j.org">Neo4J</a> <a target='_blank' href="https://gephi.org/plugins/neo4j-graph-database-support/">plugin</a> for <a target='_blank' href="http://gephi.org/">Gephi</a> was released. Gephi is an open-source <span class="highlight">visualization</span> and <span class="highlight">manipulation</span> tool that allows users to <span class="highlight">interactively browse and explore graphs</span>. The graphs themselves can be loaded through a variety of file formats. Thanks to <a target='_blank' href="http://twitter.com/#!/muzPayne">Martin Škurla</a>, it is now possible to load and lazily explore graphs that are stored in a Neo4J data store.</p>
<p style="text-align: justify;">In <a target='_blank' http://datablend.be/?p=554>one of my previous articles</a>, I explained how Neo4J and the <a target='_blank' href="http://tinkerpop.com/">Tinkerpop</a> framework can be used to load and query RDF triples. The newly released Neo4J plugin now allows to visually browse these RDF triples and perform some more fancy operations such as <span class="highlight">finding patterns</span> and <span class="highlight">executing social network analysis algorithms</span> from within Gephi itself. Tinkerpop&#8217;s <a target='_blank' https://github.com/tinkerpop/blueprints/wiki/Sail-Ouplementation>Sail Ouplementation</a> also supports the notion of <a target='_blank' href="http://www-ksl.stanford.edu/software/jtp/doc/owl-reasoning.html">RDF Schema inferencing</a>. Inferencing is the process where new (RDF) data is <span class="highlight">automatically deducted</span> from existing (RDF) data through <span class="highlight">reasoning</span>. Unfortunately, the Sail reasoner cannot easily be integrated within Gephi, as the Gephi plugin grabs a lock on the Neo4J store and no RDF data can be added, except through the plugin itself.</p>
<p style="text-align: justify;">Being able to visualize the RDF Schema reasoning process and graphically indicate which RDF triples were added manually and which RDF data was automatically inferred would be a nice to have. To implement this feature, we should be able to push graph changes from Tinkerpop and Neo4J to Gephi. Luckily, the <a target='_blank' href="https://gephi.org/plugins/graph-streaming/">Gephi graph streaming plugin</a> allows us to do just that. In the rest of this article, I will detail how to setup the required Gephi environment and how we can stream (inferred) RDF data from Neo4J to Gephi.</p>
<p>&nbsp;</p>
<h3>1. Adding the (inferred) RDF data</h3>
<p style="text-align: justify;">Let&#8217;s start by setting up the required Neo4J/Tinkerpop/Sail environment that we will use to store and infer RDF triples. The setup is similar to the one explained in <a target='_blank' http://datablend.be/?p=554>my previous Tinkerpop article</a>. However, instead of wrapping our <em>GraphSail</em> as a <em>SailRepository</em>, we will wrap it as a <em>ForwardChainingRDFSInferencer</em>. This inferencer will listen for RDF triples that are added and/or removed and will automatically execute RDF Schema inferencing, applying the rules as defined by the <a target='_blank' href="http://www.w3.org/TR/2004/REC-rdf-mt-20040210/">RDF Semantics Recommendation</a>.</p>
<script src="https://gist.github.com/1380489.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">We are now ready to add RDF triples. Let&#8217;s create a simple <span class="highlight">loop</span> that allows us to read-in RDF triples and add them to the Sail store.</p>
<script src="https://gist.github.com/1380494.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">The inference method itself is rather simple. We first start by parsing the RDF <span class="highlight">subject</span>, <span class="highlight">predicate</span> and <span class="highlight">object</span>. Next, we start a new transaction, add the statement and commit the transaction. This will not only add the RDF triple to our Neo4J store but will additionally run the RDF Schema inferencing process and automatically add the inferred RDF triples. Pretty easy!</p>
<script src="https://gist.github.com/1380511.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">But how do we retrieve the inferred RDF triples that were added through the inference process? Although the <em>ForwardChainingRDFSInferencer</em> allows us to register a listener that is able to detect changes to the graph, it does not provide the required API to distinct between the <span class="highlight">manually added or inferred RDF triples</span>. Luckily, we can still access the underlying Neo4J store and capture these graph changes by implementing the Neo4J <em>TransactionEventHandler</em> interface. After a transaction is committed, we can fetch the newly created <span class="highlight">relationships</span> (i.e. RDF triples). For each of these relationships, the <span class="highlight">start node</span> (i.e. RDF subject), <span class="highlight">end node</span> (i.e. RFD object) and <span class="highlight">relationship type</span> (i.e. RDF predicate) can be retrieved. In case a RDF triple was added through inference, the value of the boolean property <span class="highlight">&#8220;inferred&#8221;</span> is <span class="highlight">&#8220;true&#8221;</span>. We filter the relationships to the ones that are defined within our domain (as otherwise the full RDFS meta model will be visualized as well). Finally we push the relevant nodes and edges.</p>
<script src="https://gist.github.com/1380584.js"></script>
<p>&nbsp;</p>
<h3>2. Pushing the (inferred) RDF data</h3>
<p style="text-align: justify;">
The streaming plugin for Gephi allows reading and visualizing data that is send to its master server. This master server is a REST interface that is able to receive graph data through a JSON interface. The <em>PushUtility</em> used in the <em>PushTransactionEventHandler</em> is responsible for generating the <a target='_blank' href="http://wiki.gephi.org/index.php/Specification_-_GSoC_Graph_Streaming_API">required JSON edge and node data format</a> and pushing it to the Gephi master.   </p>
<script src="https://gist.github.com/1380606.js"></script>
<p>&nbsp;</p>
<h3>3. Visualizing the (inferred) RDF data</h3>
<p style="text-align: justify;">Start the Gephi Streaming Master server. This will allow Gephi to receive the (inferred) RDF triples that we send it through its REST interface. Let&#8217;s run our Java application and add the following RDF triples:</p>
<script src="https://gist.github.com/1380624.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">The first two RDF triples above state that a <span class="highlight">teacher</span> <span class="highlight">teaches</span> a <span class="highlight">student</span>. The last RDF triple states that <span class="highlight">Davy</span> <span class="highlight">teaches</span> <span class="highlight">Bob</span>. As a result, the RDF Schema inferencer deducts that <span class="highlight">Davy</span> must be a <span class="highlight">teacher</span> and that <span class="highlight">Bob</span> must be a <span class="highlight">student</span>. Let&#8217;s have a look at what Gephi visualized for us.</p>
<p><a target='_blank' href="http://datablend.be/wp-content/uploads/gephi1.jpg">
<p align="center"><img width="550" src="http://datablend.be/wp-content/uploads/gephi1.jpg" alt="gephi" /></p>
<p></a></p>
<p style="text-align: justify;">Mmm &#8230; That doesn’t really look impressive <img src='http://datablend.be/wp-includes/images/smilies/icon_smile.gif' alt=':-)' class='wp-smiley' /> . Let&#8217;s use some formatting. First apply <span class="highlight">Force Atlas lay-outing</span>. Afterwards, scale the edges and enable the labels on both the edges and the nodes. Finally, apply <span class="highlight">partitioning</span> on the edges by coloring the arrows using the <span class="highlight">inferred</span> property on the edges. We can now clearly identify the inferred RDF statements (i.e. Davy being a teacher and Bob being a student).</p>
<p><a target='_blank' href="http://datablend.be/wp-content/uploads/gephi2.jpg">
<p align="center"><img width="550" src="http://datablend.be/wp-content/uploads/gephi2.jpg" alt="gephi" /></p>
<p></a></p>
<p>&nbsp;</p>
<p style="text-align: justify;">Let&#8217;s add some additional RDF triples.</p>
<script src="https://gist.github.com/1380683.js"></script>
<p>&nbsp;</p>
<p style="text-align: justify;">Basically, these RDF triples state that both <span class="highlight">teacher</span> and <span class="highlight">student</span> are <span class="highlight">subclasses</span> of <span class="highlight">person</span>. As a result, the RDFS inferencer is able to deduct that both <span class="highlight">Davy</span> and <span class="highlight">Bob</span> must be <span class="highlight">persons</span>. The Gephi visualization is updated accordingly.</p>
<p><a target='_blank' href="http://datablend.be/wp-content/uploads/gephi3.jpg">
<p align="center"><img width="550" src="http://datablend.be/wp-content/uploads/gephi3.jpg" alt="gephi" /></p>
<p></a> </p>
<p>&nbsp;</p>
<h3>4. Conclusion</h3>
<p style="text-align: justify;">With just a few lines of code we are able to stream (inferred) RDF triples to Gephi and make use of its powerful visualization and analysis tools to explore and inspect our datasets. As always, the complete source code can be found on the <a target='_blank' href="https://github.com/datablend/blueprints-streaming-sail-inferencing">Datablend public GitHub repository</a>. Make sure to surf the internet to find some other nice Gephi streaming examples, the coolest one probably being <a target='_blank' href="http://www.youtube.com/watch?v=2guKJfvq4uI">the visualization of the Egyptian revolution on Twitter</a>. </p>
<p></p>]]></content:encoded>
			<wfw:commentRss>http://datablend.be/?feed=rss2&#038;p=260</wfw:commentRss>
		<slash:comments>0</slash:comments>
		</item>
	</channel>
</rss>
