Showing posts with label voice. Show all posts
Showing posts with label voice. Show all posts

Thursday, 17 July 2014

(Innovation) This Friendly Robot Could One Day Be Your Family’s Personal Assistant.

Jibo-inline
 Jibo
For many families, the tablet has become the central, shared computing device in the home. It’s a hub for learning, for entertainment, and for staying connected. But what if your tablet was even more interactive? What if it woke up when you came home, recognized your face, and suggested a couple of things you might want for dinner? What if, when asked a spoken question, it could tailor its answer directly to you, instead of just offering a blanket response?
A new device called Jibo can do these things, and it could mark the next step in group computer interaction in the home. But Jibo isn’t a tablet at all: It’s a robot.
Specifically, Jibo is a social robot. You talk to it, ask it questions, make requests. It talks back, provides answers, and takes care of grunt work like setting reminders or scouring the web. It’s meant to act as a helper and a partner in a variety of household experiences, much like a physical embodiment of Siri, Google Now, or any of the voice-activated concierge services available on our smartphones or tablets.
But unlike those handheld touchscreen devices, Jibo tries to act like more of a participant than a tool, as if it’s a part of the family. It has a big round head, and a face that “looks” around the room. The foot-tall, bulbous body can rotate to address the person speaking. It even leans a bit when it turns to face you, as though it’s listening more intently.
Jibo is only a prototype right now. The team behind it, headed by founder Cynthia Breazeal, who is also director of MIT Media Lab’s Personal Robots Group, hopes to bring it to market in time for the 2015 holiday season. Curious early adopters can join the crowdfunding campaign that begins today. The pre-sale price tag is $500 for early backers, and $600 for a developer kit. That’s a little more than the cost of a good tablet. And Brezeal is clear about how Jibo is designed to perform the same types of interactions families currently use tablets for, but to do so with a physical presence that fits into human lives in a more natural way than just another touchscreen.
Like a tablet, Jibo can take photos and videos. It can pull up information from the web or an app, it can act as a teleconferencing device, and it can be used to queue up books or videos. Using a mixture of facial and voice recognition (as well as an iOS and Android app), it personalizes these experiences for you. You can ask Jibo to order your favorite take-out Chinese meal after arriving home from a late night at work. Or tell it to display an e-book on its face-screen, turning a storybook into an interactive, theatrical experience for you and your child. It can recognize and greet you when you get home, or remind you to make an important phone call in between the day’s errands.
“We need technology to transcend the world of information into a more humanized realm,” Breazeal told WIRED. The connected home of the future shouldn’t feel cold and computerized, operated with Star Trek-like voice commands, she says. It should be warm and personal, interacting with us on an emotional level in addition to being able to perform useful tasks.
And thanks to the mobile computing revolution, for the first time, sensors and processors are small, efficient, and cheap enough for something like this to take the form of a robot that’s both priced and sized reasonably enough for consumers.
“Something like this is a nice bridge between devices and tablets and robots that we imagine in science fiction,” Breazeal says.

How It Works

One of Jibo’s key features is human and facial recognition. Using a stereo camera system, it can distinguish people from their background surroundings so it knows when there’s a person in the room. In particular, it can recognize faces, so it knowswhich human it’s talking to. When development is complete, Jibo will also be able to recognize facial expressions so it can guess your mood and cater its interactions to your current state of mind.
On-board hardware includes a 360 degree mic array so the robot can perform sound isolation, identifying when it’s being spoken to even if the person talking is not right next to it. Dual speakers supply its voice and other audio. Wi-Fi and Bluetooth radios keep it connected. A quad-core ARM processor act as the brains. On its face is a circular LCD touchscreen, and its plastic “skin” is also touch-responsive. A 3-axis motor system allows the top section to spin all the way around on the base. While it’s meant to stay plugged in the majority of the time, it does include a battery so you can move it around the house for short periods.

Interface and Design

Though Jibo is still a prototype, Breazeal’s team developed a demo to show what the robot will eventually be fully capable of in terms of looks and behavior. The appearance is close to final. Jibo actually looks a lot like Eve from the movie Wall-E, at least in the prototype I saw. The body is shiny, circular and white. The head is spherical, though a chunk is cleanly sliced out of it so a flat LCD display can act as its face. “For a while, we were excited about curved displays, but we realized that the technology wouldn’t be ready and robust enough,” Breazeal says.
Jibo-inline5
 Jibo turns to look at you when you talk to it.
The head and body can both rotate 360 degrees, so the robot can rotate to look at whoever is speaking to it, or just swivel and twist animatedly as it responds and interacts with you (kind of reminiscent of the Keepon robot).
As for the onscreen user interface, Breazeal added a character animator to the team to handle that task. Instead of some sort of app or list menu as an interface, or a human-like face, Jibo’s screen displays a simple, white sphere. This ball can morph into other graphical elements: a clock, an illustration of the weather, a heart, a smile. It’s designed to be dynamic and easy to read from across the room. It comes across as friendly, familiar, and expressive, all without being too cute, or verging anywhere near the uncanny valley. It’s technology humanized, but not necessarily in humanoid form.
While the prototype is expectedly rough around the edges—the LCD is low-res, and the robot’s movements are sometimes too abrupt and swift to seem natural—the potential is clear.
Jibo takes what we’ve learned from smartphone and tablet experiences, specifically from voice interactions in systems like Google Now, and builds on it. It does much of what the software on your devices can already do—learn your preferences, predict your needs—but it does everything with more personality. And whether Jibo succeeds or fails depends a lot on how that personality jibes with the humans who have to live with it.

(Fact) From Alzheimer's to ADHD: what doctors can diagnose from your voice alone



If Guillermo Cecchi wants to figure out if you've taken MDMA or meth, all he needs is a computer and a recording of your voice. Cecchi is a computer scientist at IBM, and part of a growing community of scientists who think our voices can reveal far more than our sex, age, or cultural origins. He thinks it can also unlock the mind — and the various psychological and neurological states our brains may be experiencing at any given time.

"This is exactly what psychiatrists do every day: they talk to the patients," Cecchi says, "but we used machine learning and mathematics to replicate it."

In a study published earlier this year, Cochi used recordings of short interviews to determine which drug his test subjects had been given prior to the experiment. His results rely largely on language and its meaning. "What we did on the analytics side was to use machine learning techniques that can measure things like semantic distance" — the symbolic distance between words with related meanings. "Chair" and "table" are semantically closer than "chair" and "flower" for instance. "We can identify individual interviews with high accuracy with regards to the drugs they took just by computing the semantic distance to between a handful of concepts."

PEOPLE ON ECSTASY DON’T SAY "LIKE" AND "YOU KNOW" AS OFTEN

With regards to MDMA, those concepts were friendliness, rapport, and empathy. "There was a higher similarity to these words in the interviews with a high dose of ecstasy," Cecchi says. He also found that people on MDMA used fewer "catchphrases" and jargon. When contemporaries talk to each other, "the word ‘like’ is typically 10 percent of the words." But people on ecstasy don’t use terms such as "like" and "you know" as often. Their speech, he says, is much more fluid.

Yet Cecchi’s drug-related work represents only one example of the information that our voices contain. He’s also used voice recordings to measure speech disturbances in manic depressive patients, and people who suffer from schizophrenia. Moreover, in recent years, scientists have begun to investigate the voice’s potential for diagnosing Parkinson’s disease, Alzheimer’s disease, sleepiness, depression, and even ADHD.

FROM HYPERACTIVE TO SLEEPY VOICES

Jorg Langner is a mathematician and musicologist at a Berlin-based company calledAudioProfiling. He think ADHD isn’t just about movement or ability to focus. That’s why his team is working on diagnosing children with ADHD using voice recordings. "Speech rhythm of an ADHD child" is different from a child without ADHD, he says. "The length of syllables are less equal in length." This is but one example of the measures he makes, and he says that, so far, his team has classified 1,000 previously diagnosed children with "above 90 percent" accuracy.

THE "SPEECH RHYTHM OF AN ADHD CHILD" IS DIFFERENT

Langner is also developing technology that will detect when someone is too sleepy to drive. When we’re tired, he says, our "speech rhythm isn’t so precise, it’s inexact." It’s also "not very pronounced."

Jarek Krajewksi, a psychologist at the University of Wuppertal in Germany, is working on a similar project — except his team wants to apply sleepiness detection to air traffic controllers. "Sleepiness can be detected with a classification accuracy of about 75-80 percent on unseen speaker, and 80-85 percent on known speaker" in a matter of seconds, he wrote in an email to The Verge.

But detecting sleepy air traffic controllers is just the start for Krajewski. "We have developed a depression-detection system based on 200 subjects," he said. "Another phonetic approach deals with measuring alcoholization, anxiety, confidence, leadership states or personality." He also wants to build a dataset for vocal influenza detection.

NEUROLOGICAL CLUES

Other researchers are taking a more neurological approach. "Our studies essentially looked at speech patterns in patients with Parkinson’s disease," says Rahul Shrivastav, a speech scientist at Michigan State University. "People with Parkinson’s experience changes in their voice quality, in the way they produce their sound, so vowels and consonants aren’t clear," he says. "These are very subtle, they aren’t not obvious just listening to it, but with a computer you can do much more."

Shrivastav’s team is in the early stages. So far, they’ve characterized the vocal changesthat occur when the disease is more advanced, but they hope to replicate the findings in newly diagnosed patients. This is important, he says, because there’s "no gold standard test. There’s a whole variety of symptoms that a neurologist will look at and a lot of time they will give the right drugs for Parkinson’s and if the symptoms go away, then that’s what you have." That process means that patients can go more than a decade without being diagnosed — a reality that voice diagnosis, Shrivastav hopes, will be able to change.

SOME PEOPLE CAN'T TRAVEL TO SEE A NEUROLOGIST. VOICE RECORDINGS CAN HELP

Max Little, a research fellow at MIT and the director of the Parkinson’s Voice Initiative, is also working on developing vocal diagnostic techniques for Parkinson’s. His team can obtain 99 percent accuracy in lab-based diagnostic tests, but Little notes that getting that level of accuracy isn’t "nearly as easy" with telephone-quality voice recordings. The group is now working on accurate telephone-based diagnostics. This is crucial, Little says, because many people can’t travel to a neurologist. "For them, a piece of software running on a smartphone would be perhaps the only lifeline they have to get useful information about their symptoms."

Alzheimer’s disease might also hold a future with vocal diagnostics, said Karmele Lopez de Ipiña, a computer scientist at The University of the Basque Country in Spain, in an email to The Verge. "The deterioration of spoken language immediately affects the patient’s ability to interact naturally with his or her social environment," she said, "and is usually also accompanied by alterations in emotional responses." Her team used spontaneous speech analysis to identify features, like speech fluency, to detect Alzheimer’s disease. Combined with an emotional response test, the technique boasts over 90 percent accuracy in discriminating Alzheimer’s patients from healthy controls. The ultimate goal of the research, Lopez de Ipiña said, is to identify the disease before the first clinical symptoms appear.

SUPPORTING DIAGNOSIS

The work done by these researchers differs from Cecchi’s because it relies more heavily on sounds — and the rhythms at which they’re emitted — than on language. Ultimately, however, both approaches rely on computers to analyze the connections that we make in our brains. "What our studies show is that we can measure mental states analytically without the intervention of a psychiatrist looking at the interview," Cecchi says.

ELIMINATING PSYCHOLOGISTS AND PHYSICIANS ISN’T THE OBJECTIVE

Of course, eliminating psychologists and physicians isn’t the objective. For Cecchi, the goal is to "codify" medical interviews for future use, so doctors at different hospitals in different cities, for instance, can make use of the data when a patient moves. "Psychiatrists don’t have the time to codify or measure in a way that can used by different psychiatrists," Cecchi says, adding that "we aren’t talking about therapy here, but the decision that is made or the diagnosis that’s made after an interview that happens in 30 minutes."

As for Langner and Shrivastav, both believe that their research will help strengthen previous diagnostic procedures by supplying an additional layer of objective testing. "The goal is to prevent misdiagnosis," Langner says. At the moment, a kid diagnosed with ADHD will have been tested using questionnaires and interviews with a doctor. In these instances, Langer says, a doctor’s impressions are crucial. "In many cases, these are good impressions," he says, "but it still has a great subjective component to it."

PRIVATE EXCHANGES IN FOREIGN TONGUES

Despite promising results, many challenges remain. One limitation is that some speech features are very personal and specific to an individual, Cecchi says. Another is culture. "European languages have a lot of things in common, not just language, but also culturally," he says, adding that his group has done voice analyses on people who speak Portuguese, English, and Spanish with similar results. "Now, what will happen with Chinese — we don’t know."

Langner hypothesizes that results will vary widely. "More problems occur when we go to Arabic, Farsi, and Mandarin," he says. "I think if you want to work with these languages, major adjustments will have to be done." But Shrivastav isn’t so sure. "There are differences across languages," he says, "but there are some hallmarks of certain diseases that will impact all of the layers in all the conditions, so the trick is to find those changes."

But the most worrisome aspect of this sort of research is probably the hit to privacy. "With more and more mobile phones, so much speech is being recorded and analyzed, it becomes such an easy signal to access," Shrivastav says. "I think in the next several years you will see a lot more neat things — not just for speech diagnosis." This is exactly the attitude that some critics worry about: already, researchers are working on a phone app to help doctors predict when someone with bipolar disorder might have a manic episode, so it’s possible that technology will soon be used by the public, and the government.

"AN ORWELLIAN 1984 WORLD WHERE OUR SLEEPINESS STATE IS NO LONGER PRIVATE."

"We could suffer from an Orwellian 1984 world where our sleepiness state is no longer private," Krajewski said. He thinks health and safety concerns may one day legitimize the use of this technology to monitor emotional and physical states. Someone who has a cold and is waiting for a bus might not be allowed inside the vehicle, for example. "According to a public-health regulation you will be not allowed to enter public transport — the bus door remains closed for you."

But the potential for that scenario remains years off. It’ll take a lot of time, and myriad willing participants, to unlock the information that our voices carry, Langner says. "Our brain is a giant network where everything is connected." This means that if we have problems in one location, it will have consequences in other regions of the brain, "especially in the parts that control speech projection," he says. "From that we hope that we can find traces of many other illnesses in speech sounds — but to find those solutions will be a very long process