Saturday, December 5, 2009

Outsourcing Fact Checking

Paul Spinrad guest blogs at boingboing and floats the idea of outsourcing fact checking (I'll support any proposal whatsoever that improves the fact checking process, believe me) but he adds the notion of, in essence, crowd sourcing linguistic annotation:

Now, what if these fact-checkers didn't just vet and correct the text? While they dig into the logic and accuracy of everything, as usual, they could also use some simple application to diagram the sentences and disambiguate the semantics into a machine-friendly representation. Just a little extra clicking, and they could bind all the pronouns to their antecedents, and select from a dropdown box to specify whether an instance of the string "Prince" refers to the musician Prince or to Erik Prince-- the president of XE, the company formerly known as Blackwater-- within an article that for whatever reason mentions both of them.

I have zero interest in diagramming sentences, mind you (because it's a dated and frankly messy pseudo-logical method of representing the syntax of a sentence), but there is a good idea at the core. While it's true that the web has given us greater access to large corpora, this corpora remains unstructured text. I'd like to see larger parsed corpora available (like the BNC).

With minimal training, editors and fact checkers could be utilized to mark up text with simple phrase boundaries and labels (this is a NP, this is a VP) as well as PP attachment ambiguity and co-reference, etc. There would be messiness in this approach too, but Breck Baldwin has noted that this can be done effectively (for recall, at least) and the major issue is adjudicating the error rate of a set of crowd-sourced raters (see my previous post here and Baldwin's original post here). A little sampling could adjudicate nicely.

Unsolved Problems in Linguistics

(pic from the Donders Institute)

Just discovered this page called Unsolved problems in linguistics. It's a rather incomplete list, but a start. This is the sort of topic that could easily form the core of a very interesting conference debate. Linguistics remains a wide open field with competing theories and emerging methodologies, and the big questions remain dark and murky. However, this page claims that the origin of language is the major unsolved problem. I definitely disagree. The main goal of linguistics, as I would state it, is to figure out how language works in the brain (hence, that is our major unsolved problem). From that, most other questions can be answered (btw, see The Language Guy's take down of a recent report regarding the word most here). As our understanding of the brain improves, so will our understanding of language. I don't dispute that understanding the origin of language could be of use, but it is hardly the center of the linguistics world )I realize the Derek Bickerton might disagree).

NOTE: After Googleing the phrase "Unsolved Problems in Linguistics" I found a number of other sites dedicated to the same topic, including a Wikipedia page; however there is clear plagiarism/borrowing going on somewhere as there is word for word similarity between these sites; not sure who's cutting and pasting from whom. But you need only go to one site to see the same stuff.

Friday, December 4, 2009

The Naked Vulnerability Of His Sentences


The grammatically whimsical author of Infinite Jest, the late David Foster Wallace, was, apparently, a prescriptivist. Blogger Amy McDaniel at HTMLGIANT recently posted what she claims is a "complete text of a worksheet from his class" (HT kottke) which is, basically, a grammar test which begins with the following admonition:

IF NO ONE HAS YET TAUGHT YOU HOW TO AVOID OR REPAIR CLAUSES LIKE THE FOLLOWING, YOU SHOULD, IN MY OPINION, THINK SERIOUSLY ABOUT SUING SOMEBODY, PERHAPS AS CO-PLAINTIFF WITH WHOEVER’S PAID YOUR TUITION

Feel free to take the test yourself here, or to troll the answers folks are giving.

Personal fav:

2. I’d cringe at the naked vulnerability of his sentences left wandering around without periods and the ambiguity of his uncrossed “t”s.

UPDATE: it's always nice to scoop Language Log. A day late and a dollar short, Chris Potts posts about the DFW grammar test here. Psst, my post title is wayyyyy more cleverer. thhhpppt!

UPDATE 2: More LL on DFW and his prescriptivist bent here.

UPDATE 3: Looks like scooping LL is becoming a habit for me (pats self on back).

Google Words


TechCrunch reviews Google's newish dictionary app here (Google's dictionary has been lurking around for awhile, but now it gets its own page here). I did a quick comparison of Google & Merriam Webster's entries for inappropriate and found they were remarkably different in scope. Google returns a lot more data (plus they provided links to other web definitions, which seemed to mostly be Wordnet links). I prefer Google's phonetic guide as it seems to be straight IPA (although their transcription of -pro- as pr'oʊ seems odd to me as it suggests a diphthong when I think they're just indicating rounding, but I never was much of a phoneticist, so no biggie).

I was particularly impressed to see some constructional patterns listed in Google entries (e.g., |'it' v-link ADJ to-inf| representing something like 'it is inappropriate to yell').

However, Miriam Webster still wins on historical data, minimal as it is.

Paul Reubens on Rails

kottke was in a goofy mood recently and started a twitter game whereby users come up with blends of celebrity names and online apps. Some of them are pretty good. Personal favs:
  • daniel craigslist
  • Gwyneth Paypaltrow
  • Sid Del.ico.us
  • Katrina and the (google) Waves (I'm a sucker for '80s retro)
  • Ali G(mail)
  • Bing crosby
  • Simon and Garflickr
  • Michael J FireFox
  • Ben Afflickr
  • Black IP's
See more at #webappcelebs.

Thursday, December 3, 2009

Lexical Decision Tasks

(screen grab from University of Essex demo)

UPDATE (September 14, 2017): the Uni Essex link is dead. Try this one from PsyToolkit instead.

Just found this online demo of a classic lexical decision experiment from the University of Essex here. Some images on the page don't seem to load, but the experiment runs just fine. It's a nice example of a simple psycholinguistics methodology that is commonly used in many experiments.

I'll let the good folks at Essex explain the task:

One of the key methods of investigating the processes involved in reading is the lexical decision task. Any model of reading needs to explain how a particular word can be selected from many similarly featured items, (known collectively as the neighbourhood). Neighbourhood size is a measure of the orthographic similarity between words (Coltheart et al., 1977). If a target word is orthographically similar to many words, then the target word is said to have a large neighbourhood (e.g the word sell has many neighbours such as tell, well, bell, yell and sill). A target word which is orthographically similar to few words is described as having a small neighbourhood (e.g. deny only has the neighbours defy and dent. In lexical decision tasks, Andrews (1989), found that words from large neighbourhoods elicit quicker responses than words from small neighbourhoods. This finding has been observed in a number of studies (e.g. Laxon et al., 1992: Scheerer, 1987). The facilitatory effect of neighbourhood size suggests that presentation of a target word results in activation of all the lexical entries which are similar to the target, and this local activation somehow speeds up target access. However, the precise nature of this facilitatory effect is a matter of continuing debate.

Now go enjoy the demo!

BTW, check out these other online psycholinguistics experiments here.

Wednesday, December 2, 2009

Thinking Words (part 1)

(image from make-noise.com)

I’d like to present a brief lesson in contemporary linguistic research with the goal of showing that we live in a marvelous age of quick and ready research tools freely available to even the most humble of internet users. Hence, a little effort goes a long way. My point is that when we make claims about language usage (and by "we" I mostly mean those of us who present our claims about language to the public via the interwebz) we need not make such claims based on our intuitions and emotions; rather, we can perform a little due diligence in a way that linguistic pontificators of the past simply could not. And bully for us.

My subject for today’s Full Liberman is this classic example of language mavenry from Prospect magazine: Words that think for us by Edward Skidelsky, lecturer in philosophy at Exeter University (HT Arts and Letters Daily). In this article, Skidelsky laments the following “linguistic shift”:

No words are more typical of our moral culture than “inappropriate” and “unacceptable.” They seem bland, gentle even, yet they carry the full force of official power. When you hear them, you feel that you are being tied up with little pieces of soft string. Inappropriate and unacceptable began their modern careers in the 1980s as part of the jargon of political correctness. They have more or less replaced a number of older, more exact terms: coarse, tactless, vulgar, lewd. They encompass most of what would formerly have been called “improper” or “indecent.”…“Inappropriate” and “unacceptable” are the catchwords of a moralism that dare not speak its name. They hide all measure of righteous fury behind the mask of bureaucratic neutrality. For the sake of our own humanity, we should strike them from our vocabulary.


UPDATE: A very lively discussion of the meaning of the words in question (something I largely ignore) has broken out on Language Log here)

This article makes four testable linguistic claims:
  1. The words inappropriate and unacceptable have increased in frequency over the last couple decades.
  2. This frequency increase is due to replacing other words: coarse, tactless, vulgar, lewd, improper, and indecent.
  3. These other words are “older”
  4. These other words are “more exact”
With a little investigation using entirely freely available online linguistics tools, we can easily fact check each of these claims. In the interest of time, I'll answer the first two together.

First and Second -- Has the frequency of inappropriate and unacceptable increased since the 1980s? & have they replaced the other words?

In order to quickly get some data, I took this to mean the frequency of the first two words have increased while the frequency of the other words have decreased since the 1980s (is this is an unfair interpretation?. In any case, that’s how I operationalized my methodology.). Thanks to Mark Davies excellent resource, the TIME Corpus of American English (100 million words, 1923-2006, requires registration, but it's free) we can quickly get a snapshot of the frequency of each word’s usage for the last 9 decades (not bad, huh? Thanks Mark!!).

Caveat: raw frequency is a poor data point by itself. What we really need is a way to compare apples to apples and oranges to oranges, and the problem we have is different sized corpora for each decade. Fear not, Davies does this work for us. His handy dandy interface allows us to report frequency per million, thus giving us comparable frequencies across different decades.

Using the TIME corpus, I discovered the frequency per million of each word per decade. Then I entered that data into a spread sheet. I used Excel 2007 to create a line graph of these frequencies.

Here's the relevant data:


And here's the graph:

UPDATE (2hrs after original post): original graph was confusing (same graph, just confusing labels) so I fixed it.

What this shows us is that both inappropriate and unacceptable do in fact show a rise in frequency (consistent with Skidelsky's claim), but starting in the 1960s, not 1980s. However, unacceptable shows a more recent dramatic decline, which is inconsistent with his claim. Lewd actually made a bit of a comeback in the 1990s (thank you Mr. Clinton?), but has since dropped back (it's a bit of a jumpy word, isn't it?). The other words do seem to be falling off in usage, consistent with Skidelsky's claim. So the picture is not quite what Skidelsky thinks it is, though he does seem to be on to something.

UPDATE: See myl's plot of this same data (but grouping the words as Skidelsky does) here which suggests that "'coarse', 'tactless', 'vulgar' etc. declined until WWII and then stayed about the same, perhaps with an additional decline in past decade; while 'inappropriate' and 'unacceptable' rose gradually from the 1930s to 1970 or so, and then leveled off. " The plot does suggest that we could view the two groups as having roughly inverted frequency, somewhat conforming to Skidelsky's hunch.

Third -- Are these other four words “older”?

Unfortunately, as I am no longer affiliated with a university, therefore I have no access to the OED (I’ve decided not to pay the $295 for their individual subscription. Condemn me if you must). If anyone would care to look those up and post them in comments, I’d be happy to update. Most of these words have multiple senses and the question is, when did the most relevant sense enter usage? For that, the OED is most valuable. Again, you can do that work for me, or send me a check for $295.

However, a simple search of the Merriam Webster online dictionary gives us a quick answer:

unacceptable = 15th century
inappropriate = 1804
coarse = 14th century
tactless = circa 1847
vulgar = 14th century
lewd = 14th century
improper = 15th century
indecent = circa 1587

This data suggests these five words fall into roughly two groups:

A -- words that entered the language around the 19th century
  • Set A = inappropriate, tactless
B -- words that entered the language around the 15-16 centuries
  • Set B = unacceptable, coarse, vulgar, lewd, improper, indecent
This grouping does not conform to Skidelsky’s assumption that inappropriate & unacceptable fall together in a newer class and the others in an older class.


UPDATE: much thanks to commenter panoptical who provides the following OED dates which appear to largely confirm the Merriam Webster dates, with the notable except of lewd which dates back to Old English it seems...does have a certain Beowulf ring to it, doesn't it?

unacceptable: 1483
inappropriate: 1804
coarse: 1424
tactless: 1847
vulgar: 1391
lewd: c890
improper: 1531
indecent: 1563


Fourth -- Are the other words "more exact"?

Finding a way to empirically test this is a challenge I will take up in later post (you can see Wordnet coming, can't you?). It will require teasing apart senses and relationships between senses (oh my, I wish I had the OED right now...).

TV Linguistics - Pronouncify.com and the fictional Princeton Linguistics department

 [reposted from 11/20/10] I spent Thursday night on a plane so I missed 30 Rock and the most linguistics oriented sit-com episode since ...