Monday, November 12, 2012

Art is Art: Two Galleries


I went to see Edward Tufte's gallery, ET Modern, at 547 West 20th Street. He's a sort of demigod of data visualization. Now he does art such as the above, which announces that "ART IS ART / AND EVERYTHING ELSE / IS EVERYTHING ELSE".

To get to ET Modern, I walked by the Jack Shainman Gallery at 513 West 20th Street. It looked interesting. I went in. It was showing a collection of work by Hank Willis Thomas called "What goes without saying".


Those are hand-painted blow-ups of old button designs. I explored these and other works, and then I went on to ET Modern to see Tufte's aluminum Feynman diagrams.


Tufte has made quite a lot of these aluminum Feynman diagrams. Some of them are hung like this, which brought to mind the most memorable thing I think I've ever overheard at a modern art museum:
Now this, this would go well with my sofa.

I do like Feynman diagrams quite a lot, to be sure. After I had gone to the two galleries once, I went back to "What goes without saying" to check it out again.


The second time at this gallery, I happened on a group that was being led by the artist himself, Hank Willis Thomas. It was a group of students whose professor knew the artist. I joined the group and felt welcome enough. Mr. Thomas talked about his work, something about his intent, his process, what he thought was good and interesting about various things - it was very nice. His work is not necessarily suitable for hanging over couches at country clubs. His work is communicative. It is emotional. It is meaningful. I am afraid my photo does not capture it very well. His work bears close examination. It bears consideration and reflection. At the very least, it has something to say.

I went to ET Modern once more to get a last picture. As I was leaving I asked the man who sat at a desk by the door, selling Tufte's books and things from the kind of ersatz gift shop there, why there was a big red "emergency stop" button by the exit. His explanation was a bit awkward, but he did demonstrate to me that it was completely non-functional, a kind of cartoon joke on the wall. What he said was, "It stops customers who don't buy things here."

I did think some about Tufte's work. "ART IS ART /  AND EVERYTHING ELSE / IS EVERYTHING ELSE" - I was more than a little tempted to declare that Tufte's work falls clearly on one side of that division, and that Thomas' work falls on the other. But there is still a lot to like in Tufte's work, his books, even his art.


"IT'S MORE COMPLICATED THAN THAT"

Sunday, October 28, 2012

Success in Natural Language Processing is Human-Level Intelligence

There's a lot of talk about Natural Language Processing, NLP, using computers to deal with lots of text. The current state of the art is like cooking with only a colander and whatever ingredients fall out of the tree in your back yard. People are entirely too excited about a collection of weak-sauce results.

  • Sentiment analysis is a joke on the unpopular kids (brands) desperate to know if people like them. "How many times did they say the words Nike and Love in the same sentence? Huh? How many?" AKA let's-reduce-all-human-discourse-to-one-linear-scale.
  • More general word frequency analysis can be fun, just like the index of an arbitrarily long book. And you hit problems with grammatical changes right away, so you start using some clever stemming approach, and that either makes things better or worse, and the machine is sure as heck not going to know which it is.
  • Co-occurrence? That's the best you got?
  • Okay Google's machine translation is pretty cool, but Chinese Room is not real understanding or analysis.
Take a look at this public service ad on the NYC subway:


Here's what it says:

MTA
.info
What's next?
Poetry is back
in Motion.
Many of you felt parting was not such sweet sorrow.
So we're bringing poetry back in a very artful way.
Hopefully, you'll feel transported.
Improving, non-stop.

When I was in Korea and studying Korean, I tried to read and understand signs that I came across. If nothing else, I should be able to read the signs in the subway, right? Even if you speak English as a first language, if you don't regularly ride the subway in New York, you may not know what this sign is really saying. If you're clever, you can sort of guess, but probably not perfectly.

MTA is the Metropolitan Transportation Authority. This isn't stated in the ad, but humans can probably guess that this public-service-announcement-looking sign on the subway is probably from the subway people.

"What's next?" is a rhetorical question, and in fact a sort of MTA advertising series as they tell us what updates and changes are happening now or in the near future.

"Poetry is back in Motion" refers to the MTA "Poetry in Motion" series, a separate initiative that puts short poems in subway ad slots. Apparently they stopped doing this for a while, and now they're going to be back. Or maybe they've just failed to sell all the ad slots. Who knows.

Next our friends the MTA allude to Juliet's parting words to Romeo from her balcony. Apparently by the end of the play both Poetry in Motion and all subway riders will be dead. But seriously, this allusion just doesn't make sense. Does no one understand the feeling Juliet was conveying? Honestly?

The "feel transported" is actually kind of a fun double-meaning pun. The "non-stop" is less fun but would probably be better if it wasn't following a bunch of other junk just like it.

I suppose you could say that the MTA has really done a noble job in making a dull message a little more fun. No doubt. The point is that really understanding even this fairly simple message is not so easy. What if your NLP doesn't have the NYC-subway-rider plug-in? Or the dual-meaning-pun plug-in? The Shakespeare plug-in? I suspect that until machine text analysis is done by an embodied learning computer with human-equivalent intelligence, we will be limited to the frankly unimpressive kinds of tools that we have so far. To make my suggestion even less helpful, I suspect that as soon as such technologies exist, they will have the same drawbacks that humans do. Perhaps computer users will all have to be managers. Will I have to give my computer the weekend off? Maybe I should have... ramble mode OFF

Saturday, October 27, 2012

metric-driven vs. data-driven


Bitsy Bentley, director of data visualization at GfK Custom Research, gave a good talk on Monday at Pivotal. It was mostly about visualization, but the part that resonated most with me was a related point she made about the difference between being metric-driven vs.data-driven. Here's how she illustrated where most groups currently are and the direction she thinks they need to move:

Inline image 1

Some people in the audience on Monday were confused about the distinction between metrics and data. I think it's an absolutely vital distinction, and one that hits close to home when I think about work that I'm often asked to do. A metric is a particular reduction from some subset of your data. It can be reported, rewarded, punished, used for other decisions... Metrics can certainly have value, but being focused just on some metrics is not the same thing as being data-driven. The stories in data, the real information, they frequently resist reduction to metrics - certainly to the limited collection of metrics you happen to already have. And metrics frequently obscure rather than elucidate. At best they give you a rough what - rarely a useful why or how.

As a hypothetical example, you might feel good watching a metric march in the right direction for a number of years - but if it does move in the wrong direction, that metric doesn't tell you why or what to do about it. If it was evidence of success before, is it evidence of failure now? What if you aren't doing anything differently? To really make decisions based on data requires more than just monitoring metrics. And I don't just mean you need the right metrics rather than the wrong ones - metrics are necessarily reductive, and even if you have the best metric perspectives on your data, they are still perspectives, not the data itself.

One conclusion is that it's often better to plot all of your data, as much as possible, to have a chance at understanding it before you start reducing it to numeric summaries. A corollary might be that we should question metrics that don't show the whole picture. Another conclusion is that we need to spend more time dealing with the data itself in order to understand its nature, to identify which metrics might aid understanding and which effectively stymie it, to determine what is signal and what is noise, and vitally to spend more time looking for new things than we spend recreating old things and then quickly close the loop by acting on new insight (reporting, changing policy, etc.) - and then go back to looking for the next thing.

I think this could be something to think about: are we data-driven, or are we merely metric-driven?

Ponder the divine wisdom of data cat:

Inline image 2

Sunday, August 12, 2012

Exploring Everyday Things with R and Ruby

Exploring Everyday Things with R and Ruby
Sau Sheong Chang

This book came out recently. Somebody had suggested it might be fun, so I read it. Sure enough, it is fun. It introduces a lot of really good stuff, perhaps briefly and imperfectly (and you will read "lot" instead of "plot" in at least one place) but it conveys a sense of wonder and possibility. It reminded me of the books of science experiments that I grew up with. This book is structured a bit like they were, with about eight investigations into various things. It guides you through building a digital stethoscope and processing the data it produces - and then it does the same for using a digital camera to take your pulse by reading differences in red intensity as a result of varying oxygen concentrations in your blood. It's really quite neat.


The focus is not really on building physical objects though - it's on the computer side, for simulation and analysis. The techniques weren't really new to me - and in fact in one place the author spends nearly a full page explaining the Pythagorean Theorem - but the spirit of boldly applying techniques to interesting problems is a good one. I would feel pretty good about recommending this book to a middle or high school student with an interest in technology, and any others with curiosity. To be fair, there were good pointers to things I wasn't intimately familiar with, like the Shapiro-Wilk normality test, and while not novel or very deep, the introduction to and work with ggplot2 in R is definitely widely applicable. I'm still not sure I like Ruby more than Python, but you can quickly get a feel for doing things in Ruby as well. It's a fun little book for getting you thinking, and then hopefully looking for more information and working on your own experiments.

Sunday, August 5, 2012

Memorization vs. Understanding is really Breadth vs. Depth

It is more difficult to think and communicate about things without names. Almost always, when learning something it is helpful to attach useful language. Importantly, by far the more time-consuming task is the understanding of the concept, not the learning of the term. For any given learning then, there is no real advantage to avoiding the details of terminology.

For example, students could in theory learn about evolution without learning the term "evolution" - but it would not be a better way to learn. Students could learn about a president's decisions without learning the president's name - but again, this would not be an improvement.

It is possible to memorize terms without fully understanding underlying concepts, and even if this allows a student to pass a poorly-designed exam, it is clearly not an admirable learning goal. If there are a very large number of things to learn, however, it may seem that there is only time for learning their names.

Memorizing terms alone is not a complete education. But the apparent alternative, somehow learning without learning terms or details of place and time, is not a strong alternative. We should embrace the language that accompanies learning. It may be that we need to focus on a smaller corpus to allow time for deep understanding, but that understanding will not live long without the words to discuss it.

Wednesday, June 27, 2012

Doubles are sufficient for all data relations

I was interested a while ago in RDF, which stores information in triples, like "the sky" -> "has the color" -> "blue". I think there's too much structure in the triples structure. This can all be reduced to doubles, or just a simple directed graph, through judicious use of blank nodes. So the example becomes "the sky" -> blank node that could be interpreted as having-a-property, that blank node -> "has the color", that blank node -> "blue". (If you want you could add "has the color" -> "relation" and "blue" -> "property" or similar.)

This is of course not necessarily the most efficient way to store information, or the easiest to work with, but it reduces the structure so that all the primitives are of the same type, and I don't think that the structure can be simplified any further, so in that sense it represents a kind of absolute in data structuring.

I don't know what this kind of data structure is called by others or if there is any existing work around it. I would like to know!

Thursday, June 7, 2012

How hiring works now

The problem with the job market is not that there aren't enough jobs, it's that the jobs there are can't be done by enough people. What unites them is not even that they are technology jobs; many are visual design jobs that happen to use technology. I think what really unites them is that they are habits-of-mind jobs. They are jobs for people who spontaneously do and make things - not to fill up a resume, but because they are naturally active, productive people.

Anyway, a lot of people are hiring these days. Here's a bit from a notice I saw on an email list, with my comments:

1. Do not send us a resume. Please. Don't. We won't read it.
I'm so glad other people are starting to agree with me about this. Resumes are awful.


2. Email us at ////// at /////////////.
I remember a time when teachers told me I had to go buy resume paper to print my resume on. Now most people are understanding that email is how people communicate. I'd like for companies to be required to disclose whether they have a fax machine - so I know which companies not to invest in.

3. Tell us about yourself; your hopes, dreams, desires or, better yet, how you like to code.
I like all of these, but especially the response to that last bit would tell you a lot about whether a person actually knows anything.

4. Send us examples or your work: code snippets, urls, github, etc.
This is the main point, I think: What matters is what you've done. I don't care if you aced every course at Harvard and MIT and have a recommendation from Obama. Show me what you do. I don't know that portfolios are always appropriate at every stage in education, especially because it's at least as bad as resume-focus when people focus on portfolio-stuffing.

5. Be prepared to Skype with us!
Duh.