Showing posts with label speech synthesis. Show all posts
Showing posts with label speech synthesis. Show all posts

Sunday, March 16, 2008

On The Cognitive Properties of Skin

After posting on the voiceless phone call story below, I began searching around for more information on how the device actually works. Failing to find any relevant patents pending (suspicious, I thought) I began searching for information on Michael Callahan, the wunderkind who appears to be the principle inventor, though many are probably involved.

After some searching, the most specific information I have yet found on the technology behind the voiceless phone was found in this article from the University of Illinois at Urbana-Champaign Engineering department website. Note the passage I have emphasized:

“Once we hit upon the idea of direct input, we were off and running,” explained Thomas Coleman, a project team member. The young researchers discovered that information sent from the brain can be accurately measured through the conductive properties of the skin. Typically, according to Coleman, these measurements are obtained through rigid metallic electrodes which neither respond to natural movements of the body nor to increasing skin moisture. They often become very uncomfortable under prolonged use.

"Our system uses proprietary technology to gather neurological information through encapsulated conductive gel pads, shielding the embedded electrode from the skin,” Coleman said. ‘The Audeo’ device we developed applies gentle pressure over the vocal cords, while the form-fitting band automatically adjusts in diameter, accommodating head and neck movements to maintain efficient contact.”


From there, team members created a computer program, which reads the intercepted neurological signals, and communicates a ‘response,’ both on the screen and as an audio signal. Initial work centered on determining the differences between a ‘yes’ and a ‘no’ response, which could be recognized by the computer. The software has since been enhanced to effectively ‘learn’ and adapt to the user’s neurological signals without the need of extensive training. The equipment analyzes the user during a one-time calibration process and generates a personalized user identity.
(my emphasis; quote marks had to be manually inserted to replace funny characters, but i tried to represent the original faithfully)

This is how far removed from serious neuroscience I am. I had no clue. I realized some information could be gathered from the skin, like Galvanic skin response, but I must say I’m shocked to learn that phonemic information regarding unarticulated utterances can be retrieved from the skin around a person’s neck. Clearly, there is more to this story. I’ll keep digging.

Friday, March 14, 2008

Wireless Phone Calls and Speech Production

There is a new viral video going around involving a “voiceless phone call”. Tom Simonite writes on NewScientist.com:

A neckband that translates thought into speech by picking up nerve signals has been used to demonstrate a "voiceless" phone call for the first time.

With careful training a person can send nerve signals to their vocal cords without making a sound. These signals are picked up by the neckband and relayed wirelessly to a computer that converts them into words spoken by a computerised voice.

[clip]
The system demonstrated at the TI conference can recognise only a limited set of about 150 words and phrases, says Callahan, who likens this to the early days of speech recognition software.

At the end of the year Ambient plans to release an improved version, without a vocabulary limit. Instead of recognising whole words or phrases, it should identify the individual phonemes that make up complete words.

I have no clue how this actually works (there’s an HMM in there somewhere, right?), but its implications for models of speech production ought to be significant. The folks over at Haskins Lab ought to be interested, I should think.

(HT Andrew Sullivan)

Here's the video. Cool stuff.


Wednesday, December 19, 2007

Speech-to-Text Searching

A colleague just pointed me to the new search engine EVERYZING which searches digital media audio and video (YouTube, podcasts, etc.) for your search terms using a commercially available speech-to-text engine. I’m not in love with the results, but it’s a great application for a classic computational linguistics technology.

How does EVERYZING work?
EveryZing creates a text index of the audio data from audio and video files, using the industry's leading speech-to-text technology from BBN Technologies, to enable search within the spoken words of media, not just within the metadata.

In the interests of full disclosure, though I do not work for BBN, my company does have some customers in common with them and we have utilized BBN products in service of a couple contracts (but I have not personally had any contact with BBN personnel or products).

TV Linguistics - Pronouncify.com and the fictional Princeton Linguistics department

 [reposted from 11/20/10] I spent Thursday night on a plane so I missed 30 Rock and the most linguistics oriented sit-com episode since ...