Showing posts with label Visualization. Show all posts
Showing posts with label Visualization. Show all posts

Monday, February 14, 2011

Clustering Connections with LinkedIn InMaps

Last month, LinkedIn announced a new application called InMaps which can be used to visualize a LinkedIn Network. LinkedIn’s aim is to enable its users to see what their network looks like and so better leverage their network, including identifying areas where it could be strengthened and extended.

As readers of this blog will know, data visualization is something in which we are keenly interested and so we went to try it out. Curiously, LinkedIn does not promote its labs area – or at least not that we could tell – even though there are some very interesting experimental applications in it (e.g. try out INFINITY ).

For our evaluation, we chose a relatively small network to evaluate because we were interested in exploring the representation in some depth. (Note: we have read comments from others that the software may be challenged dealing with very large networks in the 30,000+ region. D.J. Patil, Chief Scientist of LinkedIn notes the same in his comments on a posting on the FlowData blog: http://flowingdata.com/2011/01/24/explore-your-linkedin-network-visually-with-inmaps/#comment-63891 ).

It is recommended that InMaps is used with Firefox or Chrome rather than IE. Once you have reached the Labs page and selected the InMaps option, all you need to do is to permit the InMaps application to access your LinkedIn Connections. The application then processes LinkedIn’s connection-network representation and produces a diagram which is not dissimilar in style to Gephi (see previous blog posting: http://ichromatiq.blogspot.com/search/label/Gephi ) and indeed LinkedIn Maps is listed on Gephi’s own web site as a user of the Gephi toolkit (see: http://gephi.org/2011/happy-new-year/ )

Example of a LinkedIn InMap

Highly connected individuals within your network are represented with larger nodes and fonts. It is important to bear in mind, however, that the map is only representing the connectedness between the individuals to which you are connected. It is not showing the connectedness of those individuals within LinkedIn. So, for example, if you have a connection to individual A who happens to have a very large LinkedIn network but, for some reason, no one else in your network is connected to them, they will appear as a small node with a single link to you. If, on the other hand, you are connected to individual B who is connected to all the same people with whom you are connected, that individual is going to be represented as a large node.

We particularly liked the fact that the map highly interactive. Not only can you pan, zoom and mouse-over a node to get tool-tip information, but clicking on the node brings up their LinkedIn profile in the right hand sidebar. Very useful!

Most intriguing however is the clustering, represented by different colors. InMaps allows you to choose your own label for each cluster/color but gives little information as to how the clusters are derived except to say that they represent different affiliations such as previous employers, educational institutions or industries. Looking at the inMap shown here, it was clear that the dominating factor in the clustering was employment attribute and specifically company name.

Close-Up of "Misc" Cluster
The small red cluster on the immediate left of center is essentially a “misc” group. Looking at this in more detail, we noticed that connections based on professional organizations did not seem to be picked up – but that may have been either because the number of such connections was below the clustering threshold and/or the individuals concerned had not recorded the organization in their profile. We also noticed that one particular employer affiliation had not been clustered. In this case, the reason we believe is that this particular enterprise was so large that people often reference the operating division in which they work rather than the whole. Further the name of the enterprise has changed over the years. Since it would be an enormous task to keep track of all the changes – name and organizational structure – many Fortune 5000 companies go through, it might be useful to allow users to overlay the initial map with affiliations they know exist i.e. adding additional attributes.

We would have liked to have compared the representation produced by inMaps with those produced by other visualization tools: in particular NodeXL because that would have allowed us to add/modify attributes easily. Unfortunately while it is possible to export out your LinkedIn connections, you cannot access the connections between individual s in your network.

Overall, this is a very useful visualization tool, providing valuable insight into one’s professional network. It would be very interesting to overlay this with other perspectives including email traffic flow or twitter activity to give an extended picture of how one communicates and connects within the business and professional environment. More please!

Friday, December 10, 2010

Commetrix CMX Analyzer: Dynamic Social Network Visualization

Commetrix CMX Analyzer is a social network analysis platform from a German company Trilexis (www.trilexis.com) which originated in a research group at the Technical University of Berlin. (Note: the website, user interface and documentation are all in English.) What is interesting about this particular tool is its emphasis on the dynamics of social interactions over time. It achieves this through a data format that captures information about each individual link event including not only originator, destination and time but also user specified attributes which could include communication mode (email, IM, twitter), type of exchange (social, work, ecommerce), topic (e.g. keywords extracted from the subject).

Commetrix CMX Analyzer User Interface

A small subset of the Enron Email dataset –from the size and the individuals referenced we are guessing a single custodian - is provided for demonstration purposes. Part of our interest in this particular software is that we are familiar with the Enron dataset and had researched it using the social network analysis functionality of an eDiscovery system called MetaLINCS. We were curious to see what additional insights CMX Analyzer might provide.

CMX Analyzer is a desktop tool built in Java incorporating the 3D graphical capabilities of Java 3D and the Java Media Framework. Once we had obtained the license key, the application was straightforward to install and comes with a user guide. To date we have only been able to try it out on the sample data set provided as the process of creating new data sets requires end-user coding (of link attributes) followed by a data transformation process that requires as separate tool (Commetrix Producer) or the data being sent to Trilexis for processing by their systems.

Commetrix Data Preparation Process:

Commetrix is not as functionally or visually rich as some of the other tools we have investigated and reported on in previous blogs (e.g. Gephi, nodeXL). However, where it comes into its own is in the dynamic visualization of email communications over time. The MetaLINCs software we had used in the past had provided a “time-slider” but was essentially a “snapshot” approach. Commetrix has time-sliders too but also animates the traffic creating a unique perspective on what is, after all, a time-based series of events. (We should also warn readers that the resulting animations make for highly addictive viewing. We were totally captivated!) The start-end of the time period can be set, as can the intervals and speed of animation. It is also possible to run the time line backwards as well as forwards. This makes it possible to identify “hot spots” of communication activity between group subsets at particular points in time. In other types of communication e.g. twitter or facebook – we can see how this would provide valuable insight into the evolution of a topic of discussion or a social group.

Snapshot of Communications: Jan 2000

Snapshot of Communications: Dec 2000
Visually, Commetrix is more limited than some of the other packages we have used e.g. it is not possible to pan or zoom. There are options to change node size and color to represent parameters such as communications sent, communications received, number of direct contacts. Color schemes cannot be chosen directly but can be set to show selected attributes e.g. the following screenshot shows nodes color coded by the ‘function’ attribute where dark blue represents employees, pale blue represents directors, green represents traders, wholly purple circles represent managers and purple circles with yellow centers represent in-house lawyers. (Note: we found the use of full and semi colored circles to be somewhat confusing).

Colorcoding by Function

Included in Commetrix is an “egoview” option which allows you to select a particular node and investigate communication to and from that individual node. Links can be filtered to include only direct communications (a 1-step link) or communications involving two or more steps. The image below, for example, shows communications to and from Sara Shackleton. While this capability is helpful focusing down on traffic to and from a node, in the case of email communications if the data set is from only one custodian, the egoview has limited value when used outside that custodian as it will show only those communications that happen to have been referenced in emails sent to and from the primary custodian i.e. it is an imperfect sample.

Screenshot Showing Ego View - Tana Jones

Commetrix also comes with a Keyword filter. The intent is to allow the user to focus on interactions “about” the selected keywords. The interface is less obvious than some of the other areas and we confess to wondering if there was a bug until – rereading the manual – we realized that “In” didn’t mean “inbound” but include and “Out” meant exclude. Selecting the terms was also rather tedious as it meant scrolling through a long list of options. To validate the filtering, we took ‘california’ related terms and looked to see if Jeff Dasovitch was included, which he is – see screenshot below. It would be interesting to see this concept better developed with better keyword lists, more complex keyword filtering options and possibly the employment of automated topic determination techniques such as keyword clustering.

Screenshot Showing Use of Keyword Filter
Although the enron data set was provided only for demo purposes – having worked with this data, we were curious about two things: firstly how were the keywords derived (we guessed email subject but some of the keywords were email domains – indicating other metadata might have been used as well – and some phrases had been concatenated (e.g. ‘californiaattached’) or include a leading article (e.g.thenumber), or word fragments (e.g.’t’, ‘e’). Secondly, and more importantly, how were the “identities” of the individuals represented by the nodes resolved? This is always a major issue in email communications if the only information about senders and recipients is an email address. Most individuals have multiple email addresses – even within companies – and the names on email addresses may be difficult to resolve to a single individual. We raise this question because MetaLINCS included functionality that attempted to link individuals with their email accounts based not only on email address but also on communications patterns. Even then, many individuals/email accounts that a human would identify as probably being connected, could not be automatically linked. We are guessing that the identity of individuals was manually coded since the node table has a clean one-to-one mapping between individuals and a single email address.

In summary, while we think some of the other software we have used and researched offer better social network visualization options, we really liked the time-line animation Commetrix provides and believe it could be very helpful when studying the evolution of a network or communication patterns over time. While the keyword filtering option was disappointing in both the implementation and the demo dataset provided, we think it has obvious potential – particularly when analyzing large data sets of email, IM and twitter – in enabling users to focus in on only those communications “about” a particular topic. Of course, with that come all the provisos of using keywords as a substitute for “aboutness” but if it was combined with stemming, a better stop word list, and some form of thesaurus (to apply synonyms automatically) it would be very powerful.

Saturday, September 18, 2010

Analyzing Email Communications: An Ego-Centric Approach

As a quick scan through prior blogs will show, throughout this year we have been exploring the application of social network visualization software to email communications. Our interest has been two-fold: finding tools to support those working in legal and regulatory environments who need to examine large numbers of emails for answers to “Who, What, Where, When and Who Knew” kinds of questions and secondly, to see if this approach might provide behavioral psychologists with tools to identify and/or objectively measure, communications issues in workplace teams. In many workplace situations, email has become the primary communication mechanism whether through cultural factors (as with many IT teams) or because of distance (with geographically dispersed teams). At the same time communication issues are cited as one of the primary reasons why projects fail. It seemed to us that tools for analyzing the flow of email communications in a team might help identify team members who are outside the group, or who have significantly fewer interactions with key individuals in the team, thereby enabling remedial action to be taken.

Software we have looked at so far includes: Gephi – useful for large data sets – and NodeXL – useful for analyzing smaller groups of individuals with great options for customizing the appearance of the graphs e.g. color coding particular attributes or clusters and easy to use. Data feeds into both are organized basically as edge lists and node lists with Gephi requiring XML formatting and NodeXL spreadsheet or csv lists. (Note: in an email environment, a node is an individual – represented by either an email address or a name and an edge is the communication between two individuals with the volume of communications represented by a weight measure). The visualizations produced look at communication and clustering from a birds-eye view across the entire data set.

UCINET takes a somewhat different approach. UCINET is a social network analysis program developed by university researchers at the University of Kentucky and distributed by Analytic Technologies (see www.analytictech.com/ucinet/). There is a free trial version and relatively low cost options for students, researchers and single users.

Unlike NodeXL or Gephi, UCINET is not a complete visualization package but only the analytic engine. It is, however, integrated with a freeware program called NETDRAW. Since both are included in the download package, installation is straightforward. We did find in practice though that the package behaves like a set of separate tools operating on a common data set compared with the more integrated environments of NodeXL or Gephi. Another difference is that UCINET works on matrices not edge/node lists. Fortunately, it has an import function which accepts a standard edge list (e.g. person1, person2, weight) in excel format. The import function then converts this into a matrix for analysis and visualization.

Our test data set is the same as before: an anonymized set of email communications. For this investigation we started with a small subset of 368 nodes and 1223 edges.

NETDRAW visualization of entire email network


While NETDRAW is by no means as sophisticated as the graphical packages in Gephi or even NodeXL, where the UCINET/NETDRAW package came into its own is in its ability to hone in easily on a selected set of individuals. A checklist menu of nodes appears on the right hand side of the graph and altering the selections immediately redraws the graph showing only those individuals and their connections. We think this is very helpful when drilling down to investigate the interactions between a particular group of people.

Another great feature of UCINET/NETDRAW is its ability to visualize interactions from an “ego” perspective. By selecting an initial “ego”, the software identifies all the individuals in communication with the selected individual and produces a subgraph of communications between them. For example, simply selecting “Carmela Soprano” produced the following subgraph.

"Carmelo Soprano" Ego Network Graph


NETDRAW can be configured to represent the volume of communications as the size of the link:

Network Graph with Link Width Representing Communication Volume


Or with the volume shown in a link label:

Network Graph with Link Label Showing Communication Volume


UCINET offers a range of node centrality measures including Closeness, Betweenness, Degee and Eigenvector. (For information about what these measures represent, see previous blogs or go to: http://en.wikipedia.org/wiki/Betweenness_centrality#Eigenvector_centrality). Once the measures are calculated, nodes can be colorized to represent one of the selected measures. For example the nodes on the sub-graph below have been colorized to represent the value of the Indegree attribute.
It is also possible to filter based on a particular measure. The graph below shows the entire set filtered to show only nodes with high Eignvector counts (a measure of the importance of the individual in the network).

Network filtered by Eigenvector Measure (to show 'Important' individuals only)



UCINET/NETDRAW also has a number of algorithms for analyzing subgroups. For example, in the subgraph below (an “ego” network for Tom Hagen), it has identified 3 factions – represented by the three different colors: red, blue, black.

Graph identifying Factions within a Subgroup


An analysis of cliques in the entire set identified 60 separate groups shown in the graph below.

Graph showing the 60 cliques identified in the data set


What we liked about UCINET/NETDRAW is the ease with which we could explore the involvement of particular individuals in the network using the ego feature combined with the filtering and attribute based node coloring. We also liked the wide range of analysis options which included not only the standard centrality measures but also various clustering algorithms and analyses of cliques and subgroups. While more extensive documentation would have been helpful, (although we do appreciate that this was initially developed as a research tool), we did appreciate that whatever we did to it, it never crashed and managed to catch any errors gracefully.

Saturday, July 10, 2010

Analyzing Email Communications: Processed vs Unprocessed Data

In the previous post, we looked at using NodeXL to visualize communication patterns on emails that had bee preprocessed. In other words, we had run the original email file through a software tool that extracted metadata such as To, From, CC, Subject, Date Sent and stored it in a SQL database. The software we were using also extracted a Person’s Name from the email address.

For import into NodeXL, we created an edge list with the fields: PERSON_NAME1, PERSON_NAME2, CONNECTION COUNT (i.e. the number of communications between the people concerned) by simply querying the database and exporting into Excel. Later on, when we wanted to develop the visualization and look at clustering, we were able to use the database to generate a list of node attributes (e.g. Family Membership) and import that into NodeXL. For Gephi, we followed a similar process except that we output into the required XML format.

The benefits of this approach were brought home to us when we tried the Email Import feature in NodeXL. This function allows you to import network information from your personal email file into NodeXL and to configure the resulting network display. Unfortunately it is limited at present to import of personal email only – which limits its applicability. It would have been nice to have had the option to point it at some sample PSTs e.g. from the Enron data set. (And yes, we know there are workarounds to this and had the result of our test been exceptional, we might have spent time setting it up).


The import process is very simple – a click of the button if you want everything, slightly longer if you want to filter by recipient or time – and pretty quick. The resulting network retains the directionality of the email communications – which we had stripped out of our sample data. (Note: that was by choice, we could have retained it in the sample since the database captured the metadata field from which the name had been extracted).

However, we found the results of this approach not as clean or as insightful as when processed email data was used and it made us appreciate the value of preprocessing first:
(1) People’s names are almost always shorter than their Email address which makes the resulting node labels easier to work with and display.
(2) Using processed data, it is often possible to resolve multiple email addresses into the same identity. This is not a perfect science but a little text manipulation and some judicious review and editing can get you a long way. Some processing software will even support this process. With so many people holding multiple emails accounts: work and personal – this is not an insignificant issue.
(3) Processed data – because it gives you access to all the metadata – enables the network to be enriched with additional information about each node e.g. Organization, Domain. These attributes can then be used, for example, to cluster groups of nodes and provide additional insight (e.g. perhaps Operations isn’t communicating with Sales and vice versa). Attributes such as Month/Year Sent could be added to Edges.
(4) And if the metadata isn’t enough, and there is other information available that can be mapped to the individuals identified in the communications, (role maybe or demographics such as age and gender), with some minimal database work, the email network can be enriched with this information too.
(5) If the data is being imported from a database of processed email, the number of edge-pairs and nodes is known. If NodeXL is applied directly to an email file it isn’t and that means that you could very easily outstrip the capabilities of NodeXL which is designed to handle networks of a few thousand rather than tens of thousands of nodes.

Example of a Network of Email Addresses Showing Directionality with Nodes Sized and Colored by Eigenvector Centrality (i.e. Level of Importance) laid out using the Harel-Koren Method.

Sunday, July 4, 2010

Visualizing Email Communications using NodeXL

Email has become an integral part of communication in both the business and personal spheres. Given its centrality, it is surprising how few tools are generally available for analyzing it outside specialist areas such as Early Case Assessment tools within the litigation area: Xobni being a notable exception at the individual level. However, the rise of social network analysis, and the tools that support it, may change that. Graph theory is remarkably neutral as to whether it is applied to Facebook Friend networks or email communications within a Sales and Marketing division.

In a previous post, we reported on using Gephi – an open source tool for graphing social networks – to visualize email communications. In this post, we look at using NodeXL for the same purpose. We used the same email data set before – the ‘Godfather Sample’ – in which an original email data set was processed to extract the metadata (e.g. sender, recipient, date sent, subject) and subsequently anonymized using fictional names.

NodeXL is a free and open source template for Microsoft Excel 2007 that provides a range of basic network analysis and visualization features intended for use on modest-sized networks of several thousand nodes/vertices. It is targeted at non-programmers and builds upon the familiar concepts and features within Excel. Information about the network, e.g. node data and edge lists, is all contained within worksheets.


Data can be simply loaded by cutting and pasting an edge list from another Excel worksheet but there are also a wide range of other options including the ability to import network data from Twitter (Search and User networks), YouTube and Flickr and from files in GraphML, Pajek and UCINET Full Matrix DL file formats. There is also an option to import directly from an Email PST file which we will discuss a following post. In addition to the basics of an edge list, attribute information can be associated with each edge and node. In our “Godfather” email sample, we added a weighting for communication strength (i.e. the number of emails between the two individuals) to each edge and the affiliation with the Corleone family to each node.

Once an edge list has been added, the vertices/node list is automatically created and a variety of graphical representations can be produced depending on the layout option selected, (Fruchterman Riengold is the default but Harel-Koren Fast Multiscale as well as Grid, Polar, Sugiyama and Sine Wave options are also available), and by mapping data attributes to the visual properties of nodes and vertices. For example, in the graph shown below, nodes were color coded and sized with respect to the individual’s connections with the Corleone family: blue for Corleone family members, green for Corleone allies, orange for Corleone enemies and Pink for individuals with no known associations with the family.



The width of the edges/links was then set to vary in relation to the degree of communication between the two nodes i.e. the number of emails sent between the two individuals concerned.


Labels can be added to both nodes and links showing either information about the node/link or its attributes, as required.






Different graph layout options are available which may be used to generate alternative perspectives and/or easier to view graphs.

Harel-Koren Layout


Circle Layout


Because even a small network can generate a complex, dense graph, NodeXL has a wide range of options for filtering and hiding parts of the graph, the better to elucidate others. The visibility of an edge/vertex for example, can be linked to a particular attribute e.g. degree of closeness. We found the dynamic filters particularly useful for rapidly focusing on areas of interest without altering the properties of the graph themselves. For example, in the following screenshot we are showing only those links where the number of emails between the parties is greater than 40. This allows us to focus on individuals who have been emailing each other more frequently than the average.


In addition to graphical display, NodeXL can be used to calculate key network metrics including: Degree (the number of links on a node and a reflection of the number of relationships an individual has with other members of the network) with In-Degree and Out-Degree options for directed graphs, Betweenness Centrality (the extent to which a node lies between other nodes in the network and a reflection of the number of people an individual is connecting to indirectly), Closeness Centrality (a measure of the degree to which a node is near all other nodes in a network and reflects the ability of an individual to access information through the "grapevine" of network members) and Eigenvector Centrality (a measure of the importance of an individual in the network). In an analysis of email communications, these can be used to identify the degree of connectedness between individuals and their relative importance in the communication flow.

For example, in our Godfather sample, we have sized the nodes in the graph below by their Degree Centrality. While Vito Corleone is, as expected, shown to be highly connected, Ritchie Martin – an individual not thought to have business associations with the Corleone family, is shown to be more connected than supposed.

Node Sized by Degree Centrality


When we look at the same data from the perspective of betweenness, we see that Vito, Connie and Ritchie all have a high degree of indirect connections.

Nodes Sized by Betweenness Centrality


And the Eignevector Centrality measure confirms Vito Corleone's signficance in the network as well as Connie's, two "allies" - Hyman Roth and Salvatore Tessio and two individuals  Ritchie Martin.

Nodes Sized by Eigenvector Centrality


Last but not least, it is also possible to use NodeXL to visualize clusters of nodes to show or identify subgroups within a network. Clusters can be added manually or generated automatically. Manually creating clusters requires first assigning nodes to an attribute or group membership and then determining the color and shape of the nodes for each subgroup/cluster. In our GodFather example, we used “Family” affiliation to create clusters within the network but equally one could use organization/company, country, language, date etc.
"Family Affiliation" Clusters Coded by Node Color

Selected Cluster (Corleone Affiliates)

NodeXL will also generate clusters automatically using a clustering algorithm developed specifically for large scale social network analysis which works by aggregating closely interconnected groups of nodes. The results for the Godfather sample are shown below. We did not find the automated clustering helpful but this is probably a reflection of the relatively small size of the sample.

In the next post, we will look at importing email data directly into NodeXL and compare approaches based on analyzing processed vs unprocessed email data.

Larger Email Network Visualization

To download NodeXL, go to http://nodexl.codeplex.com//. We would also recommend working though the NodeXL tutorial which can be downloaded from: http://casci.umd.edu/images/4/46/NodeXL_tutorial_draft.pdf


A top level overview of social network analysis and the basic concepts behind graph metrics can be found on Wikipedia e.g. http://en.wikipedia.org/wiki/Social_network and http://en.wikipedia.org/wiki/Betweenness_centrality#Eigenvector_centrality

Monday, May 31, 2010

Visualizing Email Communications

Visualization of an email communication network

Social network analysis software enables the interactions and relationships between individuals or organizations to be modeled and visualized in a graphical format with individuals/organizations are represented as nodes and interactions/relationships as edges. In recent months, such software has been used extensively to model and analyze behavior in social spaces such as Facebook, Twitter etc. It seemed to us that it might also be useful in analyzing and understanding communication patterns in emails.

Emails are a key source of information in litigations (witness the publication of significant emails in the recent Goldman Sachs case) and are also monitored for compliance reasons in regulated industries. While most research of email data involves some form of keyword searching, there are occasions when it is important to understand who is in communication with whom: particularly if an investigation is at an early stage and may need to be broadened.

Understanding patterns of communication (as evidenced by email traffic) is also important when investigating why projects are failing or teams are not performing effectively. There is a substantial body of research that shows that communication issues are one of the primary reasons behind failing projects and dysfunctional teams. Analyzing and understanding the pattern of communications within a team or department can help business leaders and project managers identify where the breakdowns are occurring and target remedial action.

Gephi is an open source tool for visualizing networks (http://www.gephi.org/). It runs on Windows, Linux and Macs. While it will import files in a variety of formats (including CSV), the recommended format for importing data is .gexf – graph exchange XML format (see: http://gexf.net/format/index.html). GEXF is an XML based file format that is straightforward to generate once basic email metadata has been extracted and stored in a SQL database.

Some of the visualizations we were able to generate from an anonymized email set using this procedure are shown below. Gephi is very flexible in allowing for a range of different network representations and filtering so, for example, only highly connected individuals are shown. On the downside, it is still in alpha and, from our experience, not particularly robust. It crashed several times while we were attempting simple operations like adding text labels. While easy to import and export data, to use it effectively, some knowledge of the mathematics behind graphing and network analysis is helpful.

Visualization of Key Communicators in an Email Network


Use of Color to Show SubGroups within an Email Communication Network



Close-Up Showing Degree of Communication (as Line Thickness) between Participants in an Email Network

Friday, May 21, 2010

Using Tree Maps to Visualize Two Data Dimensions Simultaneously

Tree Maps (sometimes confusingly also known as Heat Maps) is a visualization technique in which data is represented as a series of rectangles whose dimensions are represented by color and size. It originated in displays of the values in a data matrix in which high values were represented by darker colored squares and low values by lighter squares.

Tree Maps are a useful visualization technique when exploring situations in which two variables interact or are interdependent in some way. For example, analyzing profitability by company size by state or the size and number of documents by custodian or document type.

In a logistics environment, when analyzing and ranking the opportunity presented by delivery to a set of zip codes, two dimensions of interest are: delivery area and delivery volume. The larger the delivery area, the longer the travel time, the greater the cost. The larger the volume the greater the profitability because within the limits of carrying capacity, the margin costs of additional deliveries are minimal. Zips can vary widely in land area and so the same delivery volume can represent a good business opportunity in one zip and not in another. Similarly a low delivery volume may be acceptable if the delivery area is very small (e.g. a building or a single block).

The tree map (or heat map) below shows the relationship between the area of a zip and delivery volume.


The land area of the zip is represented by the size of the individual squares: the larger the square, the larger the land area. The color of the squares represents the delivery volume: the darker the color, the greater the delivery volume. From a logistics business perspective; small very dark squares good, large light squares – bad. Using this visualization technique, it is very easy for sales and operational staff to identify which zips are likely to represent the better business proposition.

Several software packages are available which support such visualizations (see http://en.wikipedia.org/wiki/List_of_Treemapping_Software ). The one shown above was generated using LabEscape's Heat Map software (http://www.labescape.com/ ).

Sunday, May 16, 2010

Visualizing GeoSpatial Data

What do you do when you want to take a look at a large set of delivery address data to assess density, volume and spread? (And you don't want to spend too much time and money doing so.) The obvious first thought it to map it using Google.


It’s free, easy to use and the pins can be adapted to represent the volume of deliveries. And if you don’t want to code, you can use an excellent online service like GPS Visualizer (www.gpsvisualizer.com).

The problem is that a pin-based solution becomes too cluttered after the first couple of hundred addresses and no use at all if you have hundreds of thousands of addresses or want to get a quick overview of an entire region or state without drilling too far down to the street level.


One mapping techniques that can be used to represent density effectively is known as a Heat Map. A heat map is a graphical representation of data where the values taken by a variable in a two-dimensional map are represented as colors. The following example shows several hundred thousand delivery addresses visualized as a heat map.


Using this technique it is very easy to identify areas of heavy density. (See: http://www.heatmapapi.com/ for a free Google-based API that produces similar style visualizations).