Friday, June 6, 2014

would you like vocal fries with that?

Actual linguist Christian DiCanio debunks non-linguists' study about perceptions of fake-vocal fry (if The Onion did linguistics parodies, surely this would be it): Vocal fry doesn't harm your career prospects, but not being yourself just might.

Money quote:
...listeners judge the female speakers with vocal fry as sounding "untrustworthy", there is a good possibility that they are simply making such a judgment based on the speaker not sounding like herself. The better lesson that one might take home instead here is that one's job prospects are harmed if you try to talk (or act) like someone who you are not.
Read the full take-down here (including bonus spectrogram!)

PS: I knew Christian briefly when he was an undergrad at SUNY Buffalo. He was talented and motivated. Now he's slumming it at some shady, slacker *university* in Connecticut. Damn waste.

Tuesday, May 27, 2014

mathematical linguistics for high school students


I received the following email this weekend:

I'm a high school junior from southern California.

For our final project in AP Calculus class, I'm doing a presentation on the connection between mathematics and linguistics, and I stumbled on your blogpost "Why Linguists Should Study Math" while researching my topic.

I was wondering if you could point me towards some resources (that are relatively easy to understand) about how math is present in and affects our written and spoken language.
Some things that I am considering are:
- the occurrences of words in our language
- how grammar uses mathematical principles
- algorithms we use to construct sentences

Thanks,
M.
My [edited] response (suggestions from y'all as to better resources are much appreciated; I'll forward; I wanted to get a response out quickly because the final is presumably fast approaching):

M.,

Thanks for reaching out to me. Of course, I think you’ve chosen a good topic. There are two broad ways in which linguistics and math intersects:
  • How the human brain uses math in natural language (psycholinguistics)
  • How linguists use math to study and model languages (computational linguistics)
From your email, it appears you are mostly interested in #1. However, in contemporary linguistics, the two are fast becoming one. Most contemporary linguists use math as a tool.

Let me address your three areas of interest with respect to how the human brain might use math to process and produce language:

The occurrences of words in our language: For the most part, this means “frequency” which really means counting. Linguists love to count. We use large corpora of texts to count words and phrases. Lancaster University in the UK is a well-known corpus linguistics school. Their web page has a lot of good introductory information (although I find it a bit clunky looking).

UPDATE: I forgot to include the one item that most directly answers the basic question: frequency effects in language. Human's are very aware of how often they hear words. In some way, we count words automatically, even if it's not quite a specific count like 75, somehow we know which words, phonemes, syntactic structures we hear/read more than others. This gives rise to a variety of frequency effects in language processing. This is the clearest example of how the brain uses math for language.

For example, we recognize high frequency words much faster than low frequency words. The website for Paul Warren's book "Introducing Psycholinguistics" has an online demo for a word frequency task you can walk through to see how linguists study this.
 
What do linguists count?
  • Words: I’m sure you’ve seen word clouds like Wordle. This is composed of simple word frequency counts. One of the most enduring facts about word counts is Zipf’s Law which says “the most frequent word [in a corpus of texts] will occur approximately twice as often as the second most frequent word, three times as often as the third most frequent word, etc.” Why would this be true? Linguists have been studying this for decades.
  • Ngrams: sets of two-word, three-words, four-word strings, etc. This helps provide more context than mere single word frequencies. Have some fun playing around with Google’s Ngram Viewer if you haven’t already. Try plotting the change in frequency of “mathematical linguistics” and “corpus linguistics” (paste those two phrases into the search box with no quotes and only a comma separating them). Scholars are trying to use this to plot changes in culture. For example, take a look at this PDF.
  • Other: We also count many other things too, like parts of speech (verbs, nouns, prepositions, etc). We also count the co-occurrence of linguistics items that are not right next to each other. If you want to dig into more frequency fun, check out the more advanced tools at BYU. You can read more about how these tools help us study language here.

How grammar uses mathematical principles: One of the most commonly studied types of mathematical principle in language is statistical learning. A good example of this is transitional probabilities, which are sets of probabilities for what linguistic item might come next given a string of items (e.g., words or phonemes). For example, if you read “The author signed the _______”, you could guess what the blank word is based on the previous four words (most likely, it’s “book”).  This is based on the psycholinguistic tests called “Cloze tests”. Linguists have discovered that the brain tracks transitional probabilities for all kinds of linguistic items. In fact, this is one of the most robust areas of study in language acquisition. Linguists study how babies use transitional probabilities to learn language. For example, one of the most challenging problems is figuring out how babies learn to separate a continuous stream of audio noise coming in to their ears into separate words, without any knowledge of what words are or what they mean. One theory is that babies quickly learn transitional probabilities of sounds that tell them where one word ends and another begins. But transitional probabilities alone are not enough. For a challenge, try reviewing this PDF:

Algorithms we use to construct sentences: This is the most controversial area you’ve asked about. The fact is, we linguists don’t really know how the brain constructs sentences. As I mentioned above, there are models based on transitional probabilities like Markov models, a computer algorithm designed to make those same kinds of guesses we made about “book”. Markov models and Cloze tests are a good example of psycholinguistics and computational linguistics coming together. As a theoretical contrast to statistical models, there are rule-based models like formal grammars. These are not mathematical in a typical sense, but they are based on formal logic, which is the underlying foundation of mathematics. Linguistics is in the middle of a war between the formal grammar camp and the statistical grammar camp. There’s no consensus on which is the *correct* model of language. However, in the last decade or so, the statistical side seems to have gained the advantage. If you really want to dig in to this war, here’s a challenging read.

Additional Reading:
Linguists who count (the comments are especially engaging; your teacher might be particularly interested in the calculus vs. algebra debate that ensues).


I hope this gets you off to a good start. Please don’t hesitate to ask for clarifications or more resources (especially let me know if you need more intro level or more advanced level; I wasn’t sure if I hit the level right or not). I’m happy to be of more assistance if I can. As a smart, dedicated student, I’m sure you’re ready to dig in to ngrams and Markov models. But, as a high school junior in southern California with June fast approaching, I’m also sure you’re ready for the beach. Both are required for a healthy life of the mind.

Wednesday, May 21, 2014

Jobs for linguists - May 2014

California is awash in jobs for linguists this Spring...

Update 5/24/14: Branding and Marketing
Interbrand
NY, NY
Consultant, Verbal Identity

B.A. degree, backgrounds of interest include any verbal-focused or writing intensive field (e.g. Linguistics)
Apply Here

Text Analytics Consultant
Medallia, Inc. - Palo Alto,California
Bachelor'€™s degree
Background in Linguistics
Demonstrated interest in technology
Strong preference for a French or German native speaker
(Not visible on company website, found on LinkedIn, sign in required)
Apply Here

Linguistic Intern
Bosch Group
SF Bay Area
Responsibilities: Support the development of next-generation language products in the areas of speech and language technologies and systems. Support the administration of user studies
Qualifications: Senior undergraduate or graduate students in Applied Linguistics, or related fields
Apply Here

Analytical Linguist, Ads Human Evaluation
Google
Los Angeles, CA, USA
Product Management
Responsibilities: Direct, monitor, train, and manage the day-to-day work of temporary workers.
Design and implement tests on data and worker quality, analyzing and reporting on the results using Python, XML/CSS, HTML/JavaScript, database queries, and Google-internal technologies.
Work directly with engineers and statisticians to devise and run experiments to answer specific questions about advertising and product quality.
Minimum Qualifications: MA/MS or PhD degree in an analytical field (e.g., Linguistics, Cognitive Science, Statistics, Mathematics), or equivalent practical experience.
Experience with one or more of the following: Python or another scripting language, Java or C++, XML/HTML/CSS/JavaScript, SQL or specialized database query languages and/or specialized analysis software such as Matlab, R, SPSS, STATA, SAS, Praat, or E-Prime.
Experience working with large quantities of data.
Apply Here
And see my context here

Apple
Lexical Resource Manager
Education
M.A. or PhD in Linguistics or related field
Strong background in phonology
Apply Here

Friday, February 21, 2014

RIP Charles Fillmore

I never met Charles Fillmore, but he had a deep influence on my linguistics education. When I was a graduate student in linguistics at SUNY Buffalo we only half jokingly called it Berkeley East because half the faculty had been trained at Berkeley and the department had a *perspective* on linguistics that was undeniably colored by Berkeley theory. Charles Fillmore was a hero at SUNY Buffalo and it was hard to take a class that didn't reference his work. His work on constructions and frame semantics was the underpinning of my interest in verb classes and prepositions.

I can't offer any unique thoughts on the man, so I'll simply point to some folks around the web who have offered theirs:

A Roundup of Reactions

Paul Kay - Charles J. Fillmore
The magnitude of Fillmore’s contributions to linguistics can hardly be exaggerated

George Lakoff - He Figured Out How Framing Works
He discovered that we think, largely unconsciously, in terms of conceptual frames — mental structures that organize our thought. Further, he found that every word is mentally defined in terms of frame structures.

Dominik Lukes - Linguistics According to Fillmore
Charles J Fillmore who was a towering figure among linguists without writing a single book. In my mind, he changed the face of linguistics three times with just three articles (one of them co-authored).

UC Berkeley - Linguistics Department
He was a gifted teacher, a beloved mentor, a treasured colleague and friend, and one of the great linguists of the last half-century.

Arnold Zwicky - Chuck Fillmore
...with a link to a wonderful video he made about his career in 2012.

Friday, January 31, 2014

The SOTU and Reading Level

Evan Fleischer wrote a cheeky little bit about the reading level of the SOTU over at Esquire: Is the State of the Union Getting Dumber?

It was triggered by this graph in The Guardian:

Even emailed me and several other linguists to get some reactions. He quotes me, Ben Zimmer, and Angus B. Grieve-Smith. We generally agreed that trend noted by the graph probably had more to do with changing trends in who the speech is for, rather than any change in intelligence level.

It's a fun little read.

Tuesday, January 28, 2014

Anticipating the SOTU

In anticipation of President Obama's 2014 State Of The Union speech tonight, and the inevitable bullshit word frequency analysis to follow, I am re-posting my post from 2010's SOTU reaction, in hope that maybe, just maybe, some political pundit might be slightly less stupid than they were last year ... sigh .. here's to hope

BTW, Liberman has been on top of the SOTU story for a while now. here's his latest.

(cropped image from Huffington Post)

It has long been a grand temptation to use simple word frequency* counts to judge a person's mental state. Like Freudian Slips, there is an assumption that this will give us a glimpse into what a person "really" believes and feels, deep inside. This trend came and went within linguistics when digital corpora were first being compiled and analyzed several decades ago. Linguists quickly realized that this was, in fact, a bogus methodology when they discovered that many (most) claims or hypotheses based solely on a person's simple word frequency data were easily refuted upon deeper inspection. Nonetheless, the message of the weakness of this technique never quite reached the outside world and word counts continue to be cited, even by reputable people, as a window into the mind of an individual. Geoff Nunberg recently railed against the practice here: The I's Dont Have It.

The latest victim of this scam is one of the blogging world's most respected statisticians, Nate Silver who performed a word frequency experiment on a variety of U.S. presidential State Of The Union speeches going back to 1962 HERE. I have a lot of respect for Silver, but I believe he's off the mark on this one. Silver leads into his analysis talking about his own pleasant surprise at the fact that the speech demonstrated "an awareness of the difficult situation in which the President now finds himself." Then, he justifies his linguistic analysis by stating that "subjective evaluations of Presidential speeches are notoriously useless. So let's instead attempt something a bit more rigorous, which is a word frequency analysis..." He explains his methodology this way:

To investigate, we'll compare the President's speech to the State of the Union addresses delivered by each president since John F. Kennedy in 1962 in advance of their respective midterm elections. We'll also look at the address that Obama delivered -- not technically a State of the Union -- to the Congress in February, 2009. I've highlighted a total of about 70 buzzwords from these speeches, which are broken down into six categories. The numbers you see below reflect the number of times that each President used term in his State of the Union address.

The comparisons and analysis he reports are bogus and at least as "subjective" as his original intuition. Here's why:

Sunday, January 12, 2014

causation in verbal semantics

Causation is a major area of study within linguistic semantics. There is a thorough wiki page on the Causative that provides a good overview. Also, unsurprisingly, Beth Levin has written a nice discussion of the issues in these LSA 09 notes: Lexical Semantics of Verbs III: Causal Approaches to Lexical Semantic Representation.

To list the troubles with defining causation would fill a dissertation, so I won't bother here. Often, semanticists are interested in argument realization (see Levin's notes above). But there are deeper issues with causality that often go unaddressed. The deepest of all: what the hell is causality?

To this point, I ran across an old draft of a grad school buddy's qualifying paper on causation. It's just a draft, and it's old, but it had a nice section that tried to outline the constitutive criteria for causation*. I have since lost touch with this guy (I'll call him "BB"), but I thought this list of criteria is good food for though for anyone interested in causation. I post these as discussion points only. And if BB sees this, give me a buzz :-)

First, here's a taste of the range of causative types taken from the wiki page on Causation (don't be fooled by these English examples, the issues permeate all languages. Causation is tough):

  • The vase broke — autonomous events (non-causative).
  • The vase broke from a ball’s rolling into it — resulting-event causation.
  • A ball’s rolling into it broke the vase — causing-event causation
  • A ball broke the vase — instrument causation.
  • I broke the vase in rolling a ball into it author causation (unintended).
  • I broke the vase by rolling a ball into it  agent causation (intended) 
  • My arm broke when I fell  undergoer situation (non-causative).
  • I walked to the store  self-agentive causation.
  • I sent him to the store  caused agency (inductive causation).

BB's Nine Criteria for the treatment of causation (c. 2002)
  1. Change of state. The caused event must denote a change of state.
  2. Causers must be events. The causer A can not simply be an individual but must be an event.
  3. Argument sharing. The causing event must contain the causee in its representation.
  4. Impingement. There must be a clear indication of impingement between the causer and the causee such that the causer impinges on the causee.
  5. Occurrence condition. The caused event must occur.
  6. Co-occurrence condition. The occurrence of the caused event must be conditional with the occurrence of the causing event, that is, the caused event can only take place if the causing event takes place.
  7. Non-co-occurrence condition. The non-occurrence of the caused event must be conditional with the non-occurrence of the causing event; that is, the caused event does not take place if the causing event does not take place.
  8. Directness of causation. It must be apparent when indirect causation is allowable for causality in lexical items.
  9. Spatiotemporal equivalence. The causing event and the caused event must have an equivalent time and place.

BTW, I recall objecting to #5 "the caused event must occur" because of negative causative verbs like prevent (feel free to read my previous post on these kinds of verbs). I don't know how or if he addressed that in his final version.

* There's so much literature on causation, it would take years to review it all to see if anyone else has done such a thing at quite such a level (many authors mention criteria, but not quite as exhaustively). I wouldn't be surprised if there is a better variation out there, and I'm happy to post it if someone wants to point it out to me.

TV Linguistics - Pronouncify.com and the fictional Princeton Linguistics department

 [reposted from 11/20/10] I spent Thursday night on a plane so I missed 30 Rock and the most linguistics oriented sit-com episode since ...