Monday, July 29, 2013

Harry Potter Wrote Shakespeare's Sonnets!!!

No, of course not. That's silly.

But the recent outing of J.K. Rowling as the one true author of a crime novel published under a pseudonym was interesting not least because the software used to out her is freely available and, as it turns out, shockingly easy to use (too easy?*). You can read how Peter Millican and Patrick Juola uncovered the truth of Rowling's authorship in various places, such as:

Rowling and "Galbraith": an authorial analysis (Language Log)

You enjoy catching up to the rest of us who have actually been awake the last week or so. What I want to do is play. Much like I did with IBM's Text Analysis platform, I'm going to perform a few linguistic experiments with JGAAP over the next few weeks. The software Millican and Juola used is called the Java Graphical Authorship Attribution Program, or JGAAP. It's freely downloadable and user friendly.

I downloaded the software and opened the GUI in seconds (though the initial download site was spurious, an email to the developers quickly resolved that).

I'm running this on a modest laptop: Lenovo X100e with AMD 1.6GHz processor, 2.75 usable GB RAM, 32-bit Windows 7 OS.

First, I loaded three known authors:
  1. Shakespeare - a single text file with all plays.
  2. Christopher Marlowe - a single file from Gutenberg with most works.
  3. Francis Bacon - two text files: The Advancement of Learning and Book of Essays.
Then I loaded one "unknown author" for comparison (a single file of all Shakespeare's sonnets).

JGAAP provides very easy methods of adding all kinds of linguistic and document features to check and classifiers to use to categorize them. On my first try I chose 3 or 4 Canonicizers (normalizing the text for things like white space, punctuation, capitalization), 5 or 6 Event Drivers (ngrams, word length, POS, etc), 1 Event Culling (Most Common = 50, which I assume means to only care about the 50 most common tri-grams, word lengths, POSs), and WEKA Naive Bayes. Sadly, this failed after about 2 minutes and gave me an error message pointing me to log files. I couldn't find any log files, but I suspect I need to muck with my memory allocation for this heavy of processing.

Second, I wised up and I chose sparsely: 3 Canonicizers [normalize white space, strip punctuation, unify case], 1 event driver = Word Ngram-3, 1 Event Culling = most common events - 50, analysis method WEKA Naive Bayes Classifie).



This successfully produced results in about 2-3 minutes, though it thinks Francis Bacon wrote Shakespeare's sonnets (and really, who am I to disagree?).

This was but the first volley in a long battle, to be sure. But initial results are very promising. Dare I wonder if we are nearing that threshold moment when serious text analysis will require as many engineers as driving to the store requires mechanics?

*One could be forgiven for fearing that by hiding the serious intricacies of the mathematical classifiers and the more-art-than-science language models, JGAAP has put a weapon into the hands of children. I disagree (though not that strongly). My feeling is that JGAAP is to NLP what SPSS is to statistics. Serious statisticians probably just gasped in horror at the implications. But then again, serious drivers gasp in horror at the very idea of an automatic transmission. Technology made to fit the hands of the average is not as bad a thing as technical experts typically fear.

Let me pre-respond to one possible analogy: this is particularly salient a fear given the recent dust-up over bad neuroscience reporting (for example, read this). This is beside the point in that bad science journalism is its own special illness. It doesn't bear on the health of the underlying science.

Tuesday, June 18, 2013

Stuff to do in DC

A Twitter friend is coming to Washington DC for July and has blegged for local dives.  While there are guides aplenty for things to do in DC, there's nothing that beats a local's recommendation. So, for what it's worth, here's my list of what people coming to DC aught to take advantage of (admittedly heavy on NW).

But before I give you my recommendations, pleeeze deeer gawd!!!! Stand to the right, walk to the left on the frikkin escalators!!!

Okay...

Dive Bars
  • The Raven Grill: Tiny bar. You have to squeeze your way in. I watched one of the 2004 Bush v Kerry presidential debates here on a small black and white TV mounted in the corner. It was a partisan crowd, to say the least. (Mt. Pleasant, Columbia Heights green/yellow line).
  • Wonderland Ballroom: Isolated location. I thought I was lost the first time I tried to find it. Weird to be next to a school. But it's pretty awesome. Upstairs dance floor. (Columbia Heights green/yellow line).
  • Galaxy Hut: Honestly, I thought this place was a hipster dance club the first 100 times I walked by it and never gave it a second glance until someone told me it was for serious beer drinkers only. This book should not be judged by its cover. (Clarendon, orange line).
  • Stan's Restaurant: Lived near this place for a year and never gave it a second glance because it's buried in a basement. Turns out, it's a surprisingly awesome and friendly establishment. They pour their drinks like everyone is Hunter Thompson. Gawd help you if you ain't.(Thomas Circle, McPherson Square orange/blue line).
Not dives but worth the time
  • DC9: Small, but very fun live music venue. Most things DC run through DC9. (U Street, U street metro green/yellow line).
  • Bistro d'OC : Small French restaurant. Excellent food. The cheesy, touristy neighborhood grew up around them, don't blame them. They were there first. (Metro Center, orange, blue, red lines).
  • Black Cat: Like DC9, most things DC run through Black Cat (U Street, U street metro green/yellow line).
  • Busboys and Poets: The godfather of DC's soul. If you visit DC and fail to make your pilgrimage to Busboys and Poets, well, that's your choice, ain't it? (U Street, U street metro green/yellow line).
  • Twins Jazz: How could you not love a jazz club opened by Ethiopian twins. C'mon, man, This is what defines local flare. (U Street, U street metro green/yellow line).
  • ChurchKey : Beer lover's paradise. Temperature controlled down to the degree. A host of cask conditioned beers on tap. This is where beer poseurs go to die. Serious beer drinkers only, please. (Thomas Circle, McPherson Square orange/blue line).
  • Woolly Mammoth Theater (Archives metro, green/yellow line).
  • Warehouse Theater (Mt Vernon Square metro, green/yellow line).
  • E Street Cinema: What? A clean, well kept indie cinema in an easily accessible area? Who woulda thought?  (Metro Center, orange, blue, red lines).
  • West End Cinema: More indie cred than E Street, but also small, cramped theaters, kinda boring location, and they play the movies from a frikkin DVD. Meh. (Foggy Bottom, orange/blue lines).
  • Hike Rock Creek Park. Runs North-South along the district. There are some remarkably remote-seeming locations within this park, even though you're always dead center of DC. NYC's Central park ain't got that.
  • Arena Stage (Waterfront Metro, green line).
  • Capital Fringe Festival: I have always believed in the value of creativity for creativity's sake. We ain't ants. (various locations).
  • Eat at a "gourmet" DC Food truck. Food is awesome. Mobile food is awesome. Why should tacos own the food truck market? Do you hear me, Austin? 
General Recommendations

Walk
I'm a fan of seeing a city by the soles of your shoes and DC is a particularly walkable city. With smart phone maps and recommendation apps, DC becomes a good city to discover by foot. A comfortable pair of walking shoes are your best friend.

Capital Bike Share
For longer stays, DC's bike share program is great. They have daily, weekly, monthly and annual plans.

The Touristy Stuff
There's nothing wrong with taking advantage of touristy accommodations like bus tours because they hit the obvious highlights quickly and efficiently. Particularly with respect to The National Mall, most people have no clue just how big it is. Walking the monuments and Smithsonians is itself a monumental task that is damned tiring, especially in the hot, humid DC summer.

For the touristy stuff I highly recommend the following:
  1. The National Zoo (Woodley Park red line).
  2. The Lincoln Memorial (Foggy Bottom, orange/blue line).
  3. The Hirshhorn Museum and Sculpture Garden (on The Mall).
  4. Lunch at the cafe in the National Museum of the American Indian (on The Mall).
  5. The White House (yes yes, it really is that small).
  6. Wander around the Botanic Garden (on The Mall).
  7. You'll wait forever to get to the top of The Washington Monument; instead, go to the tower at The Old Post Office just a few blocks away. Quick, easy, and almost as great of a view.
  8. The interior courtyard at The Portrait Museum (Chinatown, green/yellow line).
Things to avoid
  • The Spy Museum - blah, always a line, expensive, cheesy and not worth it.
  • The Air and Space Museum - always packed and frankly, outdated. The phone in your hand has more impressive technology than that on display in this mothball museum.
  • Five Guys Burgers - this chain has pulled a perfect Keyser Söze. The only thing they ever did was con the world into believing they made food worth eating. They never bothered to actually make food worth eating. It's McDonalds with super sized salt. A coronary waiting to happen. Eat at a food truck, you won't be disappointing.
  • Georgetown. Basically, douchebag central.  Maybe 30 years ago there was some haute culture vibe worth observing, but now it's little more than corporate United Colors of Benetton, reality-show-cupcakes, polo-shirt-collar-flipped-up douchebag central. And I swear, if one more spandex-clad person goes jogging along the narrow sidewalks of M street shoving people out of their way as if their weekend jog somehow holds moral precedent over everyone else, Imma call a drone strike on their ass.
Personal Favorites
Everybody wants to know some local flare, the inside scoop. Here are some of my personal favorites (dictated somewhat by where I live). Some are well within walking distance of the touristy stuff
  1. Snack at Teaism, Penn Quarter (8th street NW, near the White House).
  2. Walk Roosevelt Island (Rosslyn Metro, orange/blue line).
  3. Take the tourist boat from Georgetown to Old Town in the evening (metro back).
  4. Ride a bike along the The Capital Crescent Trail that takes you from Georgetown to Bethesda (maybe 12 miles or so of relatively flat easy cycling). Have lunch, bike back. 
  5. Eat Ethiopian food. (U Street, U street metro green/yellow line).
  6. Sunday Brunch ... anywhere. DC is brunch crazy.
  7. Shoot pool at Bedrock Billiards (Adam's Morgan).
  8. Coffee and reading at Northside Social (Clarendon Metro, orange line).
  9. Artisphere (Rosslyn Metro, orange/blue line) .
  10. Join my Books and Banter book club.
  11. Bike the Anacostia Riverwalk Trail. 
  12. Lunch at Whole Foods, P Street. People watch. (Thomas Circle, McPherson Square orange/blue line).
  13. Listen to Kojo Nnamdi noon to 2 on WAMU.
  14. Smoke a cigar at Jay's Saloon (Clarendon, orange line.Talk about dives!).
  15. Sushi at Sushi Taro, 17th and P NW.
  16. Falafal at Amsterdam Falafal (Adams Morgan).
  17. Blues Alley (one of the few Georgetown places I go to regularly).
  18. Buy a copy of Street Sense for a dollar (from a badged vendor).
Leave DC
Plenty of stuff in the neighborhood.
  • Train to Baltimore, Little Italy and Fells Point (HonFest in June, but that's over).
  • Drive to Herndon for Indian food (Angeethi is one of my favs).
  • Bus to Eden Center for Vietnamese food, Falls Church.
  • Metro to The State Theatre for live music (East Falls Church metro, orange line, you can also bike in on the WOandD).
  • WOandD Trail, take a 50 mile bicycle ride into Virginia and back, 100 mile round trip).
  • Play the Sunday poker tournament at Hollywood Casino in West Virginia (1 hour drive).



Friday, May 31, 2013

Blame the linguists!

Pullum has let me down. His latest NLP lament isn’t nearly as enraging or baffling as his previous posts.

I basically agree with his points about statistical machine translation. I even agree with his overall point that contemporary NLP is mostly focused on building commercial tools, not on mimicking human language processes.

But Pullum offers no way forward. Even if you agree 100% with everything he says, re-read all four of his NLP laments (one, two, three, four) and ask yourself: What’s his solution? His plan? His proposal? His suggestion? His hint? He offers none.

I suspect one reason he offers no way forward is because he mis-analyzes the cause. He blames commercial products for distracting researchers from doing *real* NLP.

His basic complaint is that engineers haven’t built real NLP tools yet because they haven’t used real linguistics. This is like complaining that drug companies haven’t cured Alzheimer’s yet because they haven’t used real neuroscience. Uh, nope. That’s not what’s holding them back. There is a deep lack of understanding about how the brain works and that’s a hill that’s yet to be climbed. Doctors are trying to understand it, but they’re just not there yet.

He never addresses the fact that linguists have failed to provide engineers with a viable blueprint for *real* machine translation, or *real* speech recognition, or *real Q&A. Sorry, Geoff. The main thing discouraging the development of *real* NLP is the failure of linguists, not engineers. Linguists are trying to understand language, but they’re just not there yet.

Pullum and Huddleston compiled a comprehensive grammar of the English language. Does Pullum believe that work is sufficient to construct a computational grammar of English? One that would allow for question answering of the sort he yearns for? The results would surely be peppered with at least as many howlers as Google translate. If his own comprehensive grammar of English is insufficient for NLP, then what does he expect engineers to use to build *real* NLP?

It’s not that I don’t like the idea of rule-based NLP. I bloody love it. But Pullum acts like it doesn’t exist, when in fact, it does. Lingo Grammar is a good example. But even that project is not commercially viable.

One annoying side point worth repeating: Pullum repeatedly leads his reader towards a false conclusion: that Google is representative of NLP. Yes, Google is heavily invested in statistical machine translation, but there exist syntax-based translation tools that use tree structures, dependencies, known constructions, and yes even semantics. Pullum fails to tell his readers about this. In fact, most contemporary MT systems tend to be hybrids, combining some rule-based approaches with statistical approaches.

In Pullum's defense (sort of), I like big re-thinks (MIT tried a big AI re-think, though it's not clear what has come of it). But Pullum hasn't engaged in big-re-thinking. He makes zero proposals. Zero.

One bit of fisking I will add:
Machine translation is the unclimbed Everest of computational linguistics. It calls for syntactic and semantic analysis of the source language, mapping source-language meanings to target-language meanings, and generating acceptable output from the latter. If computational linguists could do all those things, they could hang up the “mission accomplished” banner.
How does translation work in the brain, Geoff? It’s not so clear exactly how bilinguals perform syntactic and semantic analysis of the source language, map source-language meanings to target-language meanings, and generate acceptable output. Contemporary psycholinguistics cannot state with a high degree of certainty whether or not bilinguals store words in their two languages together or separately, let alone explicate the path Geoff sketches out. Even if it is true that bilinguals translate the way Pullum suggests, it is also true that linguists cannot currently provide a viable blueprint of this process such that engineers could use it to build a *real* NLP machine translation system.

And that's what I have to say about that.

Friday, May 24, 2013

open the pod bay doors, Geoff

There’s man all over for you, blaming on his boots the faults of his feet.
― Samuel Beckett, Waiting for Godot

Geoffery Pullum posted a third lament about the current state of NLP: Speech Recognition vs. Language Processing. Here are his first two:

One: Why Are We Still Waiting for Natural Language Processing?
Two: Keyword Search, Plus a Little Magic.

I have responded twice.

One: Pullum thinks there are no NLP products???
Two: Pullum’s NLP Lament: More Sleight of Hand Than Fact.

My Third Response
The more I read Pullum’s three laments, the more I keep asking myself, “exactly what is Pullum complaining about and who is he aiming his complaints at?”

As far as I can tell, Pullum is complaining that commercial forces have lured researchers away from creating his dream of a human-mimicking android like 2001’s Hal 9000 or Star Trek’s Data.

This is like saying we’re still “waiting for NASA” because they failed to give us moon houses and jet packs! Is Pullum similarly unmoved by the Mars Curiosity rover? C’mon Geoff, it’s got a frikkin laser on its head!

He’s aiming his complaints at people who know nothing about linguistics or NLP (an easy audience to convince with straw men and misrepresentation).
...we are still waiting for natural language processing (NLP).
Who? Who is still waiting? I’m not waiting. I’m jumping head first into the ocean of NLP tools available right now. Who’s waiting?
...some companies run systems that enable you to hold a conversation with a machine. But that doesn’t involve NLP, i.e. syntactic and semantic analysis of sentences.
This is a rhetorical slight of hand because he is about to stack the deck and compare one petite tool to his grand Platonic ideal. Pullum continues to utilize a straw man definition of NLP that 99% of people who use the term do NOT agree with. He is wrong to insinuate that contemporary NLP cannot perform “syntactic and semantic analysis of sentences.” Of course it can. In the very least it exists in the form of POS taggers, chunkers, semantic role labeling, dependency parsers, etc. The fact that most VUI tools do not employ these extra processing components is mostly a function of optimization, not ontological failure. He dismisses this as merely "dialog design", but it's what gets products working for real consumers in the here and now. Pullum also unreasonably demotes phonetics as if it is not part of linguistics. There are many NLP tools related to speech recognition, which is where his third post goes. His punching bag for this argument is Automatic Speech Recognition.

By doing this, he creates a new straw man. What he actually describes is closer to what industry calls Voice User Interface (VUI). The distinction is non-trivial because VUI is a limited special case of ASR, not the whole kit and kaboodle. Yes, there are VUI systems which are designed to nudge users to provide responses within a limited predictable range, but there are also far more sophisticated ASR systems (like Nuance’s Dragon). These systems can produce text transcripts of voice that can then easily be ingested into any number of syntactic and semantic NLP processing tools. Ignoring them is journalistic malfeasance. Pretending they don’t exist is bonkers.
Labeling noise bursts is the goal [of VUIs], not linguistically based understanding.
This is true, but it’s not the whole picture. It’s true that VUIs are primarily trying to categorize noise bursts, but that’s the first step in the human language comprehension system too. It’s true that humans use some top-down context for predicting the likelihood of words in a continuous speech stream, but there’s plenty of bottom-up processing that is little more than “labeling noise bursts” (one of my favorite examples of this is Voice Onset Time for classifying speech segments). In focusing on this, VUIs are simply choosing one small part of the great human language puzzle to address.
Current ASR systems cannot reliably identify arbitrary sentences from continuous speech input. This is partly because such subtle contrasts are involved. The ear of an expert native speaker can detect subtle differences between detect wholly nitrate or holy night rate, but ASR systems will have trouble.
Pullum plays a little slight-of-hand trick here as he switches from talking about word segmentation to sentence breaking. These are two different tasks. Yes, human beings are very good at word segmentation and yes, ASR is mediocre, but ASR is better than he suggests and humans are not infallible word segmenters. He overstates his premise (as pointed out by his very first commenter). So, when he says that “The ear of an expert native speaker can detect subtle differences between detect wholly nitrate or holy night rate, but ASR systems will have trouble” he’s only kinda right. In fact, plenty of “expert native speakers” would have trouble segmenting those two phrases if spoken in isolation and a well trained ASR system could very well segment those phrases successfully.

Having said all that, I agree with Pullum’s underlying point that human language comprehension is mysteriously complex and intertwined with a host of non-linguistic processes like logic and memory (making "linguistic based understanding" very challenging indeed). But this is not really a fair indictment of contemporary NLP. Yes, NLPers typically narrow their focus in order to build working tools that solve one small part of the great language processing puzzle, but put those tools together in a pipeline and you can create some pretty impressive functionality.

As Pullum knows, human speech comprehension involves a complex mixture of processes and is not entirely understood by linguists even today. Understanding speech comprehension is an ongoing project in linguistics, not a finished one. Once linguists have a fully specified model of speech comprehension, I’m sure the engineers at Nuance would be happy to model it computationally. But until we linguists provide that, they’re stuck kludging a solution. If linguists are going to complain about NLP’s failures, it’s *us* we shall complain about.
...the extent to which speech as such is being processed and understood (i.e., grammatically recognized and analyzed into its meaningful parts) is essentially zero.
Zero? Really? ZERO!? Pullum was being disingenuous at best, obtuse at worst, when he wrote that. Again, his conclusion rests crucially on the straw-man comparison of one kind of limited ASR with his pie-in-the-sky fantasy of what NLP should be. This is unfair.

What Pullum refuses to tell his audience is that it is within the bounds of contemporary NLP to automatically segment a continuous human speech stream into words and then parse those words into many different grammatical and semantic categories like Parts of Speech, Subject-Verb relations, coreference, concrete nouns, verbs of motion, named entities, etc. All of this can be done by NLP tools today, right now, not in the future, by you if you have a few hours to download and learn the tools. For example, CMUSphinx Open Source Toolkit For Speech Recognition, Stanford NLP, and OpenNLP.

Pullum might complain that this NLP pipeline wouldn’t count because it wouldn’t accomplish its tasks the same way the human mind accomplishes those language tasks (and does them slower). But I repeat that it is linguists who have failed to specify exactly how the human brain accomplishes those tasks.

If Pullum is still waiting for NLP, it's because he's blaming his boots for the faults of his feet.

ADDENDUM: To be clear, I respect Geoffery Pullum quite a lot as a great linguist who has contributed (and continues to contribute) tremendous value to the field of linguistics, and language research in general. Anything I've written in my three responses to his NLP posts which might suggest otherwise is most likely the product of my uncertainty about his goals in writing these posts. I admit to feeling a bit free to employ some rhetorical flourish here and there partly because Pullum himself is quick with the lexical blade. It's fun to poke back.

Wednesday, May 22, 2013

VerbCorner - crowd sourcing and verb meaning

Josh, a postdoc at Harvard, has initiated an online game called VerbCorner in order to crowd source the study of the meaning of verbs. How often do you and I, the little people, get a chance to contribute to Harvard quality linguistic research? Well, apparently quite a lot these days. Research is for the masses!

Here's Josh's explanation
Dictionaries have existed for centuries, but scientists still haven't worked out the exact meanings for most words. At VerbCorner, we are trying to work out what verbs mean. Rather than try to work out the definition of a word all at once, we have broken the problem into a series of tasks. Each task has a fanciful backstory -- which we hope you enjoy! -- but at its heart, each task is asking about a specific component of meaning that scientists suspect makes up one of the building blocks of verb meaning.

Ultimately, we hope to probe dozens of aspects of the meaning of thousands of verbs. This is a massive project, which is why we need your help! We will be sharing the results of this project freely with scientists and the public alike, and we expect it to make a valuable contribution to linguistics, psychology, and computer science.
Being a verb meaning kinda guy myself, I'm very interested to see how this all plays out (literally and figuratively). My [defunct] dissertation was on verb semantics and Talmy's force dynamics. I'm really curious to see if Josh has included any Force Dynamics into this game.

Now, go play!

Tuesday, May 21, 2013

David Books, Word Classes, and Google Ngrams

David Brooks waxes poetic about word frequencies and the good old days in today's NYT: What Our Words Tell Us.

Update: Before reading my own most excellent original post below, here are two three well respected linguists who fisk Brooks' article as well:

John McWhorter: David Brooks' Favorite New Theory of Language Is Wrong. Money Quote:
...the faddish attempt to apply the Big Data approach to social psychology via Google’s Ngram viewer tool will shed much less light on these matters than many expect. In any language, concepts are expressed by several words and phrases at any given time, all of which morph eternally with the passage of time.

Robin Lakoff: What Our Words Don’t Tell Us. Money Quote:
It is hardly respectable scholarship to jump to the conclusion that changes in word frequency necessarily indicate changes in topics under discussion (new words may replace familiar ones but have similar meanings), and even if they do, it is very dubious – ethically questionable, you might say – to jump from there to the conclusion that these changes signify deep societal changes in the direction of moral decline, unless writers are prepared to make explicit and be prepared to defend their understanding of “morality” and “decline.” Social science is still, happily, distinguishable from theology.

Mark Liberman: Ngram morality. Money quote
David Brooks doesn't mention this ideological and temporal inconsistency in his sources. In general, as I've noted in discussions of his earlier columns, his "unparalleled ability to shape an intellectually interesting idea into the rhetorical arc of an 800-word op-ed piece" crucially depends on skillful editing — or revision — of his raw materials into a form that fits his theme.
My Original Post
Brooks cherry picks three recent Google Ngram analyses (by non linguists) and provides paper thin summaries of their findings, then concludes that America has lost is moral core. These analyses all depend crucially on the creation of word categories like “individualistic words” and “moral terms”. These are not quite synonyms*, but they require that the words in each class bear some semantic link between them. This begs the question: Are these groupings natural? Is there something psychologically real about them?

Linguists care about word classes quite a bit (computational linguists even more so). There are ways of constructing naturalistic sets of words. However, Brooks says nothing about how these studies performed their categorizations, so I thought I would post a quick review as it's important in judging the validity of the results.

Twenge et al
The first study by Twenge et al (which he doesn’t link to, but I do below) followed a scientifically reasonable path to create their word sets. They asked 53 Mechanical Turk participants to “generated words characteristic of individualism and communalism.” Then, they had a different set of 55 Mechanical Turk participants rate those words on a 7-point Likert scale. The top 20 words were then used as their search set. FYI, here are their two sets:

Individualistic
independent, individual, individually, unique, uniqueness, self, independence, oneself, soloist, identity, personalized, solo, solitary, personalize, loner, standout, single, personal, sole, and singularity
Communal
communal, community, commune, unity, communitarian, united, teamwork, team, collective, village, tribe, collectivization, group, collectivism, everyone, family, share, socialism, tribal, and union


UPDATE: For more on Twenge, commenter "unknown" helpfully suggests these Language Log posts:
Textual narcissism by Liberman
Textual narcissism, replication 2 by Liberman
It's all about who? by Liberman

Kesebir and Kesebir
Kesebir and Kesebir did 2 studies. In study one, they took ten words they found as synonyms of “virtue” in an unnamed thesaurus and searched Google’s Ngram for those words. Here are the ten: character, conscience, decency, dignity, ethics, morality, rectitude, righteousness, uprightness, and virtue.

In their second study, they constructed a set of 80 virtue words taken from websites about virtue in literature (e.g., honesty, patience, honor) then asked participants to rate each one as No = -1, Perhaps = 0, and Yes = 1. Then they took the 50 words with the highest averaged rating and search Ngrams for frequency.

Klein
Klein unapologetically gives no motivation for his word sets whatsoever. A “very casual paper” indeed.

The Problem
While I respect the attempt of the first two sets of authors to add some psychological reality to their linguistic categories, they fall for the same naïve assumption that plagued linguistics for hundreds of years: that people's conscious judgement of meta-linguistics is valid. For example, syntacticians discovered the folly of grammaticality judgments. I have been involved recently in a number of Mechanical Turk ratings tasks and we're finding that it is very difficult to get consistent ratings. I believe the same issue is at play here. Plus, ratings can easily be affected by context like surrounding text, yet none is given in these tasks. It's not clear what it means to rate isolated words. Word semantics by their very nature are contextual.

UPDATE: Commenter Arjan rightly brought up the great acceptability debates. One could claim that I am unfairly dismissing grammaticality judgments. And one could claim that I am not. The good folks at MIT's Tedlab have posted a few excellent resources on multiple sides of the controversy. Look under the 2010 heading on this page.

Words are not thought. These studies seem to be a variation on the “No word for X” syndrome (see here for a recent rant). Certain types of words may be used more or less frequently over some time-scale (like one century), but that doesn’t necessarily mean that we are thinking differently over that time-scale.

Unlike Brooks, I’ll link to the actual papers (all free, but the second two require email registration):

Increases in Individualistic Words and Phrases in American Books, 1960–2008. Jean M. Twenge, W. Keith Campbell and Brittany Gentile

The Cultural Salience of Moral Character and Virtue Declined in Twentieth Century America. Kesebir and Kesebir

Ngrams of the Great Transformations. Daniel B. Klein

UPDATE: *Rumor has it that WordNet has copyrighted the term "synset", so I'm being careful to avoid their cease and desist letter. Anyone know if there's truth to this rumor?

Saturday, May 18, 2013

Book Reviews

A quick (very self-serving) link fest. Here are the cognitive linguistics related book reviews I've written:

1. Adam's Tongue: How Humans Made Language, How Language Made Humans. By Derek Bickerton.

2. Louder Than Words: The New Science of How the Mind Makes Meaning. By Benjamin K. Bergen.

3. Through the Language Glass: Why the World Looks Different in Other Languages. By Guy Deutscher.

TV Linguistics - Pronouncify.com and the fictional Princeton Linguistics department

 [reposted from 11/20/10] I spent Thursday night on a plane so I missed 30 Rock and the most linguistics oriented sit-com episode since ...