Speech Recognition Headed to the Masses
From the Original Pages
Click a page to enlarge · Alt-click to open the full issue
A Profile of Lernout & Hauspie
When it comes to a forming a more natural human-computer interface, even pen computing comes second to possibly the most human of all methods of interaction: speech. The goal of true speaker-independent speech recognition, capable of accurately functioning in a range of conditions and environments, has challenged scores of researchers the world over.
Undaunted, several universities and companies have kept working on the problem and, bolstered by faster and more specialized hardware (such as RISC processors and DSPs), less expensive memory, and newer software algorithms (including neural networks), have begun to offer products and technologies worthy of serious consideration.
One such company is Lernout & Hauspie, one of the leaders in the emerging field of speech recognition. Since accurate speech recognition promises to so radically redefine how we may interact with future mobile computers, we’ve listed Lernout & Hauspie as a Company to Watch.
“Accurate speech recognition promises to radically redefine how we may interact with future mobile computers.”
Corporate Background
Headquartered in Belgium and Massachusetts, Lernout & Hauspie was founded in 1987 when it exclusively acquired speech technologies developed at leading universities in Belgium. The company currently employs over 250 people and has offices in the Europe, Asia, and the United States.
The current management team includes Jo Lernout, co-founder and co-chairman, along with a person who should be familiar to most mobile industry watchers: Gaston Bastiaens, president of L&H and former executive in charge of the Newton Systems Group at Apple Computer.
Other members of the management team include Pol Hauspie, co-founder and co-chairman managing director; Nico Willaert, vice chairman and managing director; Robert Kutnick, chief technology officer and senior vice president; and Ellen Spooren, vice president, corporate communications.
Key Products and Technologies
L&H develops and markets three main technologies. These are:
- Speech recognition
- Text-to-speech
- Speech and music coding
Reflecting its European origins, each of the technologies are available in several languages and are adaptable to multiple hardware platforms including DSPs [Digital Signal Processors] from Texas Instruments, AT&T, Motorola, and others, as well as x86-compatible and RISC processors.
In the category of recognition, L&H offers the Automatic Speech Recognition (ASR) system which performs continuous recognition using a small to medium-sized vocabulary. The system is also capable of recognizing isolated words and performing keyboard spotting.
As long as the words are relatively distinct (i.e. are not confusingly similar in pronunciation), several hundred words can be active in a vocabulary at any one time. As a specific example, a test vocabulary of the 500 most common names in New England has achieved greater than 95% recognition accuracy across a sample of six speakers, according to the company.
The system does this by performing four tasks:
- Data acquisition and preprocessing,
- Feature extraction,
- Acoustic matching, and
- Dynamic programming.
The first step involves preprocessing the signal to remove echos, increase gain, and basically clean the signal for later stages.
With feature extraction, the actual recognition process begins. This starts with a frequency analysis from which acoustic and phonetic information can be distilled. This stage also functions to suppress environmental noise and other undesirable information within the signal.
The next step involves performing an acoustic match between a series of feature vectors created earlier, and a set of acoustic-phonetic units available to the recognizer.
The final stage uses a technique known as dynamic programming to offer the best match for words and sentences based not only on the acoustic scores, but also on constraints imposed by the lexical and syntactic information incorporated into the system.
L&H’s Text-to-Speech (TTS) products operate in the opposite direction: converting computer-readable information into human-sounding speech. This technology is based on three components:
- The Linguistic Module
- The Phonetic Module
- The Acoustic Module
Each module does pretty much what its name implies: the linguistic module performs a series of lexical and syntactical analysis to create a phonetic representation of the text; the phonetic module works to add a human sounding pattern—including segmentation and duration—to the sounds; and the acoustic module converts this information into actual speech signals.
L&H’s third product concentrates on speech and music coding which offers the ability to compress these high-quantity data streams into a more efficient format suitable for storing and transmitting. A typical application would include compressing incoming messages on an all-digital telephone answering machine.
Recent Developments
While the company has made a number of announcements during the past year, we’ll concentrate on ones related to computing and consumer electronics. Most recently, Lernout & Hauspie announced that, together with Hitachi, the companies will port the Text-To-Speech and Automatic Speech Recognition products to Hitachi’s SH-3 RISC processor.
As you may recall, the SH-3 is at the heart of a number of the new Windows CE-based handhelds including those from Casio, Compaq, LG Electronics, Hewlett-Packard, and of course Hitachi. The SH-3 ranks as the most popular RISC processor, used in countless embedded applications including, but not limited to, consumer and home electronics.
Hitachi envisions a number of applications including having computers read electronic mail out loud, as well as using the speech recognition features to locate appointments and contacts on devices that have limited-size keyboards. While neither feature is expected in first generation Windows CE devices, the stage has been set for subsequent versions.
Perhaps even more notable, L&H announced at Fall COMDEX that Microsoft has also licensed the ASR and TTS products for possible inclusion into future products.
“Microsoft has licensed the company’s ASR and TTS products for possible inclusion into future products”
Craig Mundie, Senior Vice President of the Consumer Platform Division at Microsoft (which encompasses the Windows CE effort) optimistically noted that the technology should provide a strong foundation for future speech-related applications.
Mobile Business Implications
Accurate user-independent speech recognition technology has already moved out of the world of science fiction and is poised, over the next several years, to become an integral part of the next great leap forward in user interface design. The implications are substantial for both mobile and non-mobile devices and applications.
In principle, hands-free computer operation offers a number of advantages including:
- Enabling the computer user to concentrate on the task at hand, whether this involves operating a vehicle or performing some equally hazardous task.
- Allowing the computer to be used in less environmentally friendly environments, or in situations where gloves must be worn. (Mobile computers without keyboards can be more effectively sealed and protected against environmental hazards, or embedded into other devices.)
- Permitting higher volume data entry without specialized training (learning to type, for example).
- Allowing people to multitask more efficiently by providing a way for them to interact with the mobile computer while performing some other mundane task, such as actively searching for an item in inventory.
Speech recognition has broad appeal and will therefore likely appear in both consumer and industry-oriented devices within the next year. However, true speaker-independent, context-free, recognition is still likely to overtax all but the most powerful processors available today, including those that appear in today’s handhelds.
Nevertheless, the company notes that a RISC processor is well suited for the dynamic programming module described earlier when working on a limited vocabulary. In addition, L&H’s ASR is being used in commercial products based on inexpensive processors such as the Analog Devices AD2105.
In terms of microphone requirements, the L&H signal preprocessing module is designed to normalize the effect of microphones. At a minimum, a microphone should cover the frequency range of speech (i.e. at least up to 4kHz bandwidth).
The L&H ASR engine accepts PCM data sampled at 8kHz and 11kHz—the input can come from far talk microphones, close talk, or even cellular phones.
Company spokesperson Audre Pobre added: “In one commercial navigation system using the L&H ASR, a relatively inexpensive microphone is used in a far talk location (placed on the visor), providing accurate recognition even in a moving car. For high noise industrial applications, noise cancelling or directional microphones can also be used to improve accuracy, if necessary.”
“Interestingly, a reasonable candidate for speech recognition is Apple’s new Newton MessagePad 2000.”
Interestingly, a reasonable handheld candidate for speech recognition is Apple’s new Newton MessagePad 2000. Powered by a fast RISC processor (the Digital StrongARM running at 160MHz), and featuring a built-in microphone and speaker, the MP2000 seems almost ideally suited for the task.
Apple has informally confirmed that third-party developers are indeed working on the problem and that solutions should be forthcoming, possibly this year. Understanding that speech is even more difficult that handwriting recognition, the company is proceeding carefully in its promotion of possible features and technologies.
Pobre concurred with the fundamental notion: “Speech is a natural adjunct to handhelds. Between the screen and keyboard, the need is there. But two things need to happen: speech input and/or output hardware, and suitable processors.”
“The speech channels are here—all HPCs have speech output, and the Philips HPC and Newton have microphone input. As far as processors are concerned, we can run ASR, TTS, and coding very easily) on an SH-3, Motorola 68000, AD 21xx, TI C20x, x86, etc. These can be found in existing handhelds and phones.”
Transcribed from Pen-Based Computing, Volume 7, Number 1 — January 1997. Pages 12, 13, 14.