Pages

Showing posts with label junk viz. Show all posts
Showing posts with label junk viz. Show all posts

Tuesday, October 29, 2013

Junk Viz - When More is Less

There are examples of junk visualizations, and then there are examples of junk charts that just take your breath away.

The Indian news portal, FirstPost.in, which describes itself as a "trusted guide to the crush of news and ideas around you", published a story titled Shivraj set for massive victory in Madhya Pradesh: Survey | Firstpost, which has this chart (link to the image) - take a minute to study it. Then study it again. It is no optical illusion or card-trick being played here.

Friday, October 25, 2013

HBR and Junk Charts

Even the best can get it wrong. Only sometimes though, one hopes.
The venerable Harvard Business Review gets data visualizations horribly wrong. They have a post on Facebook where they contrast the costs of cancer treatment in the US and India.

The cost in the US is $22,000 on average, while in India it is shown as $2,900 (I would dispute this figure, as it looks very low).
The ratios is 7.58:1 (22000÷2900)

Wednesday, July 10, 2013

Lying with Charts - Global Warming Graph

Global warming is a serious yet controversial enough topic without bringing in bad data visualizations practices into it. The Wonkblog on the Washington Post has an article titled, "You can’t deny global warming after seeing this graph". The post reproduces a chart prepared by the World Meteorological Association that plots global temperatures by decade. While the data shows that the last decade, 2001-2010, was the hottest on record, the graph uses a broken Y-axis that begins at 13.4°C instead of starting at zero. The chart does not hide this fact, and you can see that the chart's Y-axis starts at 13.4°C, but the most visually prominent piece in the graph is, well, the graph! And it screams the message that global temperatures are going off the charts - it's time to panic. There is no denying that we as a world need to get serious about investing in alternative and renewable sources of energy like solar, wind, and even nuclear, but this graph is just plain bad.

Tuesday, April 17, 2012

Lying with Charts - Google Finance and Apple

Here's a short post on visualizations and distortions, unintentional but still there.

There was an article on the web that remarked on the rather steep fall-off in Apple's stock over the past week or so. I went over to Google Finance to take a look. What I found was interesting. I took some screenshots and have added them to this blog post.

I wanted to find out how much the stock had actually fallen, which is easily done, and how much the line chart was portraying as the fall in the stock, also fairly easily done.
Let us do some math now. Simple math, the kind I like, the only kind I can probably do now.

First, let us calculate how much the stock has actually fallen. The Google Finance page on Apple,  http://www.google.com/finance?q=aapl , tells us that on April 16, the stock closed at $580.13 - which we will round off to $580. Next, we find that its 52-week trading high was $644.
So, you can see that the stock has fallen $64 from its peak, which translates into a 10.2% fall from its peak (64/640).
So, the first number of significance is 10.2% - we will format it bold to make it noticeable.

Next, take a look at the first chart. Even with a linear scale, the problem is that the axis does NOT begin from zero. Notice the first number on the vertical axis is $420 - Google is using a broken axis, which is useful for highlighting the magnitude of changes, as in this graph, but misleading because of its very nature; it inaccurately magnifies increases and decreases. By how much? Let's calculate.

If you were to take a measure and see what is the height of the stock chart from the base to its maximum, i.e. $644, you would find it measures 4" from top to bottom - approximately.
Next, you measure the fall from the peak of $644 to the current trough of $580. It measures approximately 0.95".
So, in this chart, a peak of $644 equates to 4".
A drop of $64 measures 0.95"
Therefore, the chart plots the drop as a 23.75% drop as seen on the chart - we will bold it to make it noticeable.

There you have it - an actual drop of 10.2% looks, note, looks, like a 23.75% drop.
To put that in perspective, had the stock actually fallen by 23.75%, it would have sunk by $152. Yes, and it would have been trading at $492.

Apple Stock Graph in 2012

Even when you change the time-scale to 5 years, it does not help completely, because the vertical axis is STILL a broken axis. The inaccuracy as displayed on the chart is a lot less, but it is still there.
Apple Stock Price over 5 Years

It is only the 10-year plot that has a true, non-distorted picture of the stock. But because of the 10-year plot, the recent rise and steep 10% fall is not very visible. If you zoom only into the current year, 2012, then the distortions creep right back into the graph.
Apple Stock Price over 10 Years

What Google Finance needs to do is add an option, a checkbox, in their Settings panel to allow a user to select whether they want an unbroken axis or not - i.e., to let the charting engine plot a broken axis when it sees fit, or to always display an unbroken axis.

Saturday, June 18, 2011

Designing With the Mind In Mind-Book Review

Designing with the Mind in Mind: Simple Guide to Understanding User Interface Design Rules
Designing with the Mind in Mind: Simple Guide to Understanding User Interface Design Rules (Kindle Edition)

This is another excellent book to add to your shelf, after giving it a careful read, alongside other excellent books on information visualization.
I claim an interest in the field of data visualizations. And not just the Lego blocks in colorful arrangements type of visualizations, though that is not to gainsay their utility or ability to entertain. This interest in meaningful information and data visualizations goes back at least 8 years, to 2003, when I first started working as the product manager for visualizations in the Discoverer product from Oracle (sort of tautological - working for Oracle would presuppose that the products I worked on would be Oracle products too...). Starting in 2004 my interest in visualizations took a more detailed turn when I starting haranguing people about the utility of having interactive visualizations. Some of what I have written in my capacity as a product manager for data visualizations in Oracle BI since 2006 has made its way into the product, much more is making its way into the product, and there is much that will eventually, I hope, make its way into the product.

GUI Bloopers 2.0, Second Edition: Common User Interface Design Don'ts and Dos (Interactive Technologies)Therefore, it is but natural that I also have an interest in literature on data visualizations. To that end I have read some books and papers and blogs on the topic over the years, including Information Dashboard Design: The Effective Visual Communication of Data, The Visual Display of Quantitative Information, Envisioning Information, A Tour through the Visualization Zoo - ACM Queue, and more... There is always the humorous yet educational blog Junk Charts. Then there is the often acerbic yet valuable blog by Stephen Few (whose post first led me to this book), Visual Business Intelligence. And so on...

This year, the first book I have read on the topic, Designing with the Mind in Mind: Simple Guide to Understanding User Interface Design Rules by Jeff Johnson, is not really a book on data visualizations per-se. This book will not tell you the utility of a bar graph versus a line graph. It will not tell you what decorations to apply or not apply to graphs, whether 3D effects look good on a graph (they don't), what chart junk is (see Edward Tufte's books for that), etc... The author, Jeff Johnson, is an authority in this field, has been active in the field of HCI (Human Computer Interactions) for more than 30 years, and has worked at Xerox, Sun Microsystems, HP Labs, etc...

Don't Make Me Think: A Common Sense Approach to Web Usability, 2nd EditionHis latest work is more a book about the theory of how the mind perceives information, of how humans understand what they read, and how our eyes are attuned to paying attention to not just what's happening in front of us but also at the periphery of our vision. This is a "design" book - "Design rules often describe goals rather than actions. They are purposefully very general to make them broadly applicable, but that means that their exact meaning and their applicability to specific design situations is open to interpretation.". It is a book that informs us how some of the perceptual hard-wiring in our brains has evolved because of very sound reasons, and why information systems that tend to ignore or force their way against these perceptual conduits often fail. That you have more a vast proliferation of interfaces that are designed so as to violate these fundamental precepts of cognition is an indication of how far we still have to go in this field.

In every book on user interface design, whether specific or general, you will find the usual suspects - the Gestalt principles: Proximity, Similarity, Continuity, Closure, Symmetry, Figure/Ground, and Common. The author says these provide a useful basis for guidelines for graphic and user interface design.
... Several Gestalt principles describe our visual system’s tendency to resolve ambiguity or fill in missing data in such a way as to perceive whole objects. The first such principle, the principle of Continuity, states that our visual perception is biased to perceive continuous forms rather than disconnected segments.
Slider controls are a user-interface example of the Continuity principle. We see a slider as depicting a single range controlled by a handle that appears somewhere on the slider, not as two separate ranges separated by the handle.

A recommended practice, after designing a display, is to view it with each of the Gestalt principles in mind—Proximity, Similarity, Continuity, Closure, Symmetry, Figure/Ground, and Common Fate—to see if the design suggests any relationships between elements that you do not intend.


Information Visualization, Second Edition: Perception for Design (Interactive Technologies)If we are to read and understand what's written, like instructions on a screen, or tooltips, or the like, then it stands to reason that the typeface be easy-to-read. But the author goes beyond that, deeper, into the roots of how we read and understand, and how, therefore, poorly designed interfaces can interrupt the process by which we understand what we understand.
In other words, the most efficient way to read is via context-free, bottom-up, feature-driven processes that are well learned to the point of being automatic. Context-driven reading is today considered mainly a backup method that, although it operates in parallel with feature-based reading, is only relevant when feature-driven reading is difficult or is insufficiently automatic.
... reading can be disrupted by hard-to-read scripts and typefaces. Bottom-up, context-free, automatic reading is based on recognition of letters and words from their visual features. Therefore, a typeface with difficult-to-recognize feature and shapes will be hard to read.


Visual noise in and around text can disrupt recognition of features, characters, and words and therefore drop reading out of automatic feature-based mode into a more conscious and context-based mode.
The same goes for colors too. Color patch size and separation for example are used by our visual system to make out one color from another.
Color patch size: The smaller or thinner objects are, the harder it is to distinguish their colors
Separation: The more separated color patches are, the more difficult it is to distinguish their colors...

Color patches in chart legends should be large to help people distinguish the colors


“change blindness” (Wikipedia link) is what sometimes causes people to not pay attention to not pay attention to a message of possible importance flashed by an application. Therefore, "Don’t require people to remember system status or what they have done, because their attention is focused on their primary goal and progress toward it."

Envisioning InformationThere have been at least three books I have read this year that have ended up talking about the concept, history, and neurology of memory (Moonwalking with Einstein: The Art and Science of Remembering Everything, Talent Is Overrated: What Really Separates World-Class Performers from Everybody Else, and The Shallows: What the Internet Is Doing to Our Brains), and this book is a fourth one, if you add to the list books that cover only peripherally the topic of memory. In this book, the author describes the workings of memory, how they are formed, and what implications it has when it comes to aiding users in remember previously performed actions in a graphical user-interface.

memories, like perceptions, consist of patterns of activation of large sets of neurons. Related memories correspond to overlapping patterns of activated neurons.
...

However, it has many weaknesses: it is error-prone, impressionist, free-associative, idiosyncratic, retroactively alterable, and easily biased by a variety of factors at the time of recording or of retrieval.
...
One implication of this pattern is that interactive systems should indicate what users have done versus what they have not yet done.
...
A new face stimulates a pattern of neural activity that has not been activated before, so no sense of recognition results. Of course, a new face may be so similar to a face we have seen that it triggers a misrecognition, or it may be just similar enough that the neural pattern it activates triggers a familiar pattern, causing a feeling that the new face reminds us of someone we know.
...
In contrast, recall is long-term memory reactivating old neural patterns without immediate similar perceptual input.
...
Whatever the evolutionary reasons, our brain did not evolve to recall facts.
...
Because people are bad at recall, they develop methods and technologies to help them remember facts and procedures
...
The relative ease with which we can recognize things rather than recall them is the basis of the graphical user interface (GUI)
...
The relative ease with which we can recognize things rather than recall them is the basis of the graphical user interface (GUI) (Johnson et al., 1989). The GUI is based on two well-known user interface design rules:
  • See and choose is easier than recall and type.
  • Use pictures where possible to convey function.

And what does the author mean here?
Even insects, mollusks, and worms, without even an old brain—just a few neuron clusters—can learn from experience. However, only creatures with a cortex or brain structures serving similar functions[2] can learn from the experiences of others.
...
caveat is that some birds can learn from watching other birds.


The mind just races with the possibilities. A student peering over the shoulder of another at an exam is sure learning from the experience of others, on a lighter note.


The Visual Display of Quantitative InformationAs with other tasks, consistency within an application's interface is critical. This is also one of the primary tasks of a user interface design engineer - to ensure that different screens, different parts of an application all have the same vocabulary of interface and action. Different parts of an application are worked upon by different engineers, and this can often enough cause those parts of an application to look inconsistent in how they look and feel (the classic problem that LAF standards seek to minimize). Even with the benefit of guidelines and look-and-feel standards that are in place at most large software development companies, it is inevitable that inconsistencies can creep into the UI of an application. This is where the importance of a user-interface and user-experience design team cannot be stressed enough.
To reduce the time it takes for people to master your application, Web site, or appliance, so that using it becomes automatic or nearly so, don’t force them to learn a whole new vocabulary
....
Same name, same thing; different name, different thing. (FormsThatWork.com) This means that terms and concepts should map strictly 1:1. Never use different terms for the same concept, or the same term for different concepts. Even terms that are ambiguous in the real world should mean only one thing in the system. Otherwise, the system will be harder to learn and remember.

Performance and the perception of responsiveness are different beasts altogether, related only by the often contentious thread of individual experiences. Personal experiences can differ widely. What one considers slow is considered acceptable by someone else. In my life and times as a product manager, there have been several occasions where discussions about performance, the expectation of performance, and what can be considered as responsiveness on the part of an application and what should be considered as 'slow' have ranged from the pleasant, the cordial, to the contentious even.
Responsiveness is related to performance, but it is different. Performance is measured in terms of computations per unit of time. Responsiveness is measured in terms of compliance with human time requirements and, as described above, user satisfaction.
...
Time lag between a visual event and our full perception of it: 100 milliseconds (0.1 seconds)
...
Our brain compensates by extrapolating the position of moving objects by 0.1 second. Therefore, as a rabbit runs across your visual field, you see it where your brain estimates it is now, not where it was 0.1 second ago
To be perceived by users as responsive, interactive software must follow these guidelines:
Acknowledge user actions instantly, even if returning the answer will take time; preserve users’ perception of cause and effect
Let users know when the software is busy and when it isn’t
...
Animate movement smoothly and clearly • Allow users to abort (cancel) lengthy operations they don’t want
...
Interactive systems should avoid lengthy gaps in on their side of the conversation. Otherwise, the human user will wonder what is happening. Systems have about 1 second to either do what the user asked or indicate how long it will take.
...
It is true that meeting those deadlines on the Web is difficult—often impossible. However, it is also true that those deadlines are psychological time constants, wired into us by millions of years of evolution, governing our perception of responsiveness.


Information Dashboard Design: The Effective Visual Communication of DataEvery book on memory and cognition will also talk about the two kinds of memory that exist. One is the long-term memory, which consists of the things we remember for a long time, often as long as our lives. The other is short-term or working memory, to which is the attributed the magic number of seven, plus or minus two, which is the average number of objects a person can hold in their working memory. It turns out that while this number may not appear to be impressively high, in reality it is even lower!
This breaking down of tasks into subtasks ends with small subtasks that can be completed without a break in concentration, with the subgoal and all necessary information either held in working memory or directly perceivable in the environment. These bottom-level subtasks are called “unit tasks” (Card et al., 1983).
...
Unit tasks have been observed in activities as diverse as editing documents, entering checkbook transactions, designing electronic circuits, and maneuvering fighter jet planes in dogfights, and they always last somewhere in the range of 6 – 30 seconds.
 In conclusion, this is not the book to pick up in the middle of a time-sensitive project to get guidance on user-interface doubts. No. The time to pick this book and go through is before. Or in-between deadline-driven assignments.





Kindle Excerpt:



Tuesday, September 28, 2010

Data Visualizations - Show Some Hide Some

The Search Engine Land site had a post on Jun 29 2009 - Google: We’re Not Really That Big But If We Are, We Aren’t Bad - where charts were used; specifically, charts that Google has used to argue that while it may appear to be a big company, it is not that big when compared to some of its peers, or in relation to the size of the market itself.

Here is the data table used below (from the site):


And two pie-charts:


Let's start with the Google pie chart, which tries to highlight the diminutiveness of Google's market share.
Data visualization experts have lamented the use of pie charts in visualizing data. I will not belabor the point. Instead, I believe the same data could have been displayed more effectively using either of the two charts below

The first one is a simple vertical bar chart, while the second one is a stacked percentage vertical-bar chart. Since we are working with percentages, either chart is conveying the same information.

The data in question for Google is called out by the use of a different color, and the size of the data is made clear by the height of the 'Offline' bar. Even in comparison to the 'Other Offline' bar, the Google bar's size pales in comparison.
Adding a data label may seem redundant, but if the chart does not support hover tooltips, then adding the label, with the precise value of the series, helps.
OR


http://www.businessinsider.com/chart-of-the-day-google-is-not-that-big-after-all-2009-7 mentions the post, and has an accompanying chart:

Here the attempt is to compare each company on two variables - revenue and employees. A dual-Y-axis bar as the one used above is not good, not good at all, for such a presentation of data. What exactly is the point of plotting revenue and employees as bars in this graph?A line-bar maybe, but even that is sub-optimal IMO.

If you have to do this type of a comparison, the scatter plot is most effective since it allows for a meaningful comparison across companies too. To make this interesting, you could also add a third variable, and plot the data as a bubble chart. You could also use a butterfly graph.

When displayed as a vertical bar, what is obvious is that Google has much lower revenue, and even fewer employees when compared to any of the companies it is being compared with. Fair enough. But what the chart does display, but not tell you very clearly, is something very interesting, that I explain with charts below.

Take revenue and employees, the two measures in the chart above. To perform a meaningful comparison, it is first necessary to normalize these values first. One way is by using ratios instead. i.e., the ratio of the company's revenue to the number of employees. Divide the revenue by the employees, plot **that** data instead, and this is what you get:

Innaresting, wouldn't you say?
Google's per-employee-revenue is more than one million dollars (say it Dr Evil style and it doesn't sound as sinister, maybe funnier), while IBM is less than a fourth as much at $254k per-employee, and even Microsoft is only $600k per-employee.

What this tells us, as long a we are comparing these companies, is that Google is able to eke out a lot more revenue from its employees than other companies. Unless its employees are super, super-freaks, it can mean, among other things, that Google has achieved economies of scale far beyond its peers, like Microsoft, IBM, AT &T, and Verizon. In some ways it can also be an apples to oranges comparison, because the businesses that Google and ATT and Verizon and IBM are not exactly comparable. But, this is the set of companies that accompanies the post, and is also the set of companies that Google chose. So there.

If you were to do a similar comparison with Microsoft, taking only its Windows and Office divisions, that have a near-monopoly market share in their respective segments, I am sure you would come up with near-Google numbers, or maybe even better. Conjecture, but a fascinating one, IMO.

If you accept Google's proposition that the other companies in the comparison are in similar businesses as Google, or that they are competitors to Google, then you also have to accept the proposition that revenue-per-employee figures also should be similar. If they are not, it could be because the companies are diverisified into areas where such economies of scale do not apply; which is a fair argument to make, or that these companies are just not able to extract as much money as Google is.

Let's use another ratio. This time, I use market cap-to-number-of-employees as the ratio.

Why use this ratio? What does this tell us, if anything? Well, for one, market caps are fairly fickle numbers, and can be misleading. But since the data is from the Google table, let's use it anyway. What it may tell us is how valuable each employee is to Google's shareholders.

Simplistically speaking, each Google employee adds $5 million to the company's market capitalization.
That is more than twice Microsoft's.
That is more than sixteen times IBM's.
 
Again, assuming Google's employees are not all hyper-efficient Einsteins, and some or even many would argue that is indeed the case, it means at the very least that Google's hold on the business and industry it operates in is a lot more efficient and powerful than its competitors. Which could result from, among other things, near-monopoly pricing power.

 Let us use a third set of metrics, and this time, let's also plot a bubble chart, that can display three measures reasonably well.

What is plotted on the x-axis is market cap as a multiple of revenue. i.e. for Microsoft, the bubble in blue, this would be 184 billion (its market cap) divided by 60 billion (its revenues), to yield a figure of 3.07. And similarly for the others.
What is plotted on the y-axis is market cap as a multiple of operating profits.
The size of the bubble is the company's revenues.

What does this chart tell us?
Firstly, that Google plots as an outlier. Good outlier or bad outlier? Well, that depends on whether you are Google or its competitor.
It also tells us that Google's price-earnings ratio is way out of whack with its competitors. It could also mean that Google is incredibly over-priced, or that it has such a strangle-hold on its business that giant gains in market share, and consequently revenues and margins, are almost guaranteed over the coming years, which is why the market has driven up its market cap to such heights.

From Google's perspective, unless it intends using the currency of its market cap to make big-ticket acquisitions, such a high market cap is not really that good. It only attracts market envy and unwanted regulatory attention.

Anyway, another example of how data can be used to show some and hide some.

Wednesday, February 17, 2010

Junk Viz - the 100 slice pie chart

As junk visualizations go, while there has been enough that has been written on the drawbacks with the pie chart, this example below, from 10 Ways to Archive Your Tweets, brings out the problem in a most, shall we say, visual manner.

As far as gleaning any information from this chart goes, it's a lost cause. You would need to possess incredibly powers of being able to precisely position your mouse over a particular slice to see its value. If, on the other hand, you decided to go down the legend, see a name, and then go to the pie chart to figure out the value of the slice, there are so many colors in use that the entire exercise would be reduced to an almost futile case of trial and error.

So what in its place? A simple table would have sufficed. With perhaps an underlying data bar on the cells. You could conditionally format the table so as to color code the rows in deciles, or quartiles.

Or you could slice the data by deciles, and place a dropdown above the table so that only ten rows at a time were displayed.

Or you could display only the top 10 and the bottom 10 tweeters by default, and provide the user with an option to expand the middle 80 rows of the table.

Or almost anything else. But this pie chart.

Wednesday, November 25, 2009

Visualizations - The Pie Chart

The Telecom Regulatory Authority of India (TRAI) put up a press release, Date: November 21, 2009 Press Release:Telecom subscribers growth for the month of October 2009. ,  that has information on the telecom subscription data in India for the month of October 2009. Apart from the quite amazing piece of news that India added 16.67 million (that is 16,670,000) new wireless subscribers, and that the total telephone subscriber base now stands at 525.65 million (that is more than half a billion), the notable thing as far as this blog post is concerned is that depressing use of visualizations in the note.


A few things are obvious at first glance:
- It is a pie chart with a 3D effect.
- This is an Excel generated chart.
- There is redundancy in the chart: the slice labels contain the operator name, and then the legend at the bottom repeats the same information.
- The data is not sorted, so even if you could somehow compare these 3D slices, you would have a tough time finding which is the largest slice, which is the second largest slice, and so on.
- To find the largest slice, you are better off simply comparing the numbers. Which makes the chart itself quite unnecessary.
- The color scheme is very Excel-ish, which is to say, quite unpleasing to the eye. Excel 2007 is an improvement, for sure.
- There are black borders around the slices, which do not make the chart any better.
How to improve this?
Here are some examples:

Example 1:
You cannot really go wrong with a bar chart. This bar chart displays the same data, except now as a bar chart. Straight off you can tell from a visual inspection that "Tata" added the most subscribers, close to 25% of the net additions in October 2009.



Example 2:
I have now added data labels at the top of each bar. This makes it possible to see the precise values for each operator.




Example 3:
By now, it is clear that sorting the bars would make the data a lot more easily digestable. So what insights are now possible with this example? For one, that Reliance and Aircel and even Idea are two operators that added almost the same number of subscribers. Not very obvious from the above examples. Aircel is a relatively new operator, but seems to be growing quite fast, thanks to its aggressive advertising.




Second Chart:



This table above shows "Category wise Net Additions during the Month of October 2009'.
Notwithstanding the fact that the data here would be a lot more easy to understand if it had been formatted with commas, let us see how it may be visualized as a chart:

This chart does one thing well. It gives a sense of the difference in scale between the wireline and wireless segments. The wireless segment is growing by millions, in every circle, while the wireline segment is in decline. The decline is however minuscule. And without labels, it is difficult to gauge even the approximate values.

So, if I plot this now as a percent stacked bar chart, it looks like an improvement. What I have done is added labels to each stack. I can now see that the Metro segment showed a rise, while the other three segments showed a decline in the wireless segments.
However, this chart is sort of misleading, because it makes the wireline and wireless segments appear equal. Which, as we saw, is most certainly not the case.



As the third example, I have now plotted the same data as a stacked vertical bar chart. Not as a percent stacked chart, but simply taken the absolute values and stacked them.

The vertical chart brings out quite nicely the difference in magnitude between the wireline and wireless segments.
A problem existed for this chart also. Which is that the categories for the wireline segment are so small, that the individual stacks are barely visible, even on a chart as tall as this one. So, I have added data labels, and then manually moved the labels so that they don't overlap.