Wednesday, May 30, 2007

Microsoft Surface

After five years of keeping the project shrouded in secrecy, Microsoft today revealed its plans for Microsoft Surface, the first product in a category the company calls "surface computing." The technology, formerly code-named Milan, lets Microsoft turn a seemingly ordinary surface, such as a tabletop or a wall, into a computer. Introduced today at the D: All Things Digital conference in Carlsbad, California, Microsoft Surface is a "multi-touch" tabletop computer that interacts with users through touch on multiple points on the screen.

The concept is simple: Users interact with the computer completely by touch, on a surface other than a standard screen. "It will feel like Minority Report," promises Pete Thompson, general manager of Microsoft's surface computing group. "Very futuristic--but it will be here this year."

The product unveiled today will be Microsoft branded and available to the company's four partners--Harrah's Entertainment, International Game Technologies, Starwood Hotels, and T-Mobile--in November. Starwood Hotels plans to put Microsoft Surface devices in common areas, to provide functions such as a virtual concierge; T-Mobile will use them to enhance the cell phone shopping experience. Microsoft expects to deploy dozens of units with each of its partners by year's end.
Microsoft Surface couples standard PC components with the cameras and projectors necessary to enable surface computing. The demo unit employed a 3-GHz Pentium 4 CPU, 2GB of RAM, and an off-the-shelf graphics card with standard drivers (and Microsoft's own application layer to allow the GPU to help with sensing touch).

The images the PC outputs are displayed on the tabletop surface through a short-throw DLP projector contained inside the table; the lens is just 21 inches from the surface. The rear-projection system produces a 30-inch-diagonal, 4:3-aspect-ratio image at a resolution of 1024 by 768 at 60 Hz.

The table also houses a power supply, stereo speakers, an infrared illuminator, and five overlapping cameras that sense movement on its surface. The cameras feed images of objects on the surface--be they fingers or tagged objects such as game pieces, a Wi-Fi camera, or a digital audio player--back into the computer, where they're processed mostly in the GPU, according to Nigel Keam, one of Microsoft's architects behind Surface.

The specially treated surface's multi-touch capability has no implicit limit, says Keam. "We optimize it for 52 [points of touch], based on the most extreme reasonable scenario we could come up with: Four people with all fingers down, and 12 game pieces in the center."

One of the hardest things about working with the technology was to get the touch surface right. Developers had to walk a fine line in creating a surface that's opaque enough to hold a rear-projected image but translucent enough for cameras to see through it. "You need a strong diffuser on the topmost surface," Keam notes, "but the camera wants to see straight through the diffuser to what's on the surface. So it's a balancing act. We had to research a lot of different ways to make the surface look right, feel right, and be tough. Everything meets at this one layer."

The device's infrared capability means you can do more than just use your fingers on the tabletop surface. Tags on a Wi-Fi camera or a digital audio player, for example, could be used to transfer images, music, or playlists. Or perhaps a card could store your account information and let any Microsoft Surface unit grab your images from a central server. Tagged pieces might generate special effects for drawings or images, and puzzle pieces could act as props in interactive games.

Monday, April 16, 2007

Google, Microsoft look beyond mobile search for voice interaction

Organizing all of the world's information is no good unless that information can be accessed, and Google's recent move into free 411 searches shows that the company is serious about voice search. But Google's plans appear to extend far beyond finding local businesses. In a patent issued last year, the company outlined a full-blown natural language voice search platform, and Microsoft's own recent acquisition of TellMe indicates that a voice search war could be brewing.

The Google patent, issued in April 2006 and naming Sergey Brin as an inventor, describes a system for providing search results from voice queries, but it goes well beyond the system currently implemented as GOOG-411. In the patent, the inventors recognize the problems inherent in this sort of searching: lack of context and very short queries. "There is very little repetition in queries, providing little information that could be used to guide the speech recognizer," they note. "In other speech recognition applications, the recognizer can use context, such as a dialogue history, to set up certain expectations and guide the recognition. Voice search queries lack such context."

Nevertheless, Google believes that it can solve the problem, and the fruits of that effort appeared to be on display in GOOG-411, which can take moderately constrained input data (a business or business category, constrained by city) and return useful results. Though limited to business search at the moment, the system can handle a huge variety of accents and ways of requesting the same information, and pairing voice queries with the eventual search results is an excellent way to build up a database of useful voice snippets. This is exactly what Tim O'Reilly thinks is going on: Google is using an early implementation of the system to gather enough data to improve its voice recognition algorithms before a broader launch.

Voice search has been heating up in the last few months; one of Microsoft's largest buyouts this decade was made earlier this year, when the company acquired TellMe, which develops voice-recognition applications for use over the telephone. One of its most successful products is an automated 411 system; could Google's launch of a similar service less than a month after the Microsoft acquisition be pure coincidence? It certainly could, but it could just as easily be a show of strength from Google, which wants to make clear that Microsoft's expensive purchase can already be generally replicated by Google engineers.

New patent hints at Apple TV 2.0

A new Apple patent may give some foresight to the company's plans for offering a "true" multimedia center experience for the living room, as either enhancements to the Apple TV or an entirely new device. The patent is for a "Multi-media center for computing systems" and describes a system involving a central multimedia hardware hub that can make use of a number of external "modules," all controlled through a centralized menu system on the device.

These external modules can consist of anything that a multimedia center might want to read from, be that a computer, iPod, external DVD player, or hard drive. However, the module could also (speculatively) be something like a recording device or something to stream from/with, like a Slingbox. According to the patent, the modules would not interact with each other at all but would be managed dynamically by the central multimedia center through a main user interface. The device will also be able to function as a traffic manager, with the ability to "instantiate and keep track of media-modules, route events associated with user input, and control what is displayed."

The patent text also says that the system could take inputs from a variety of different things, such as a keyboard and mouse, over the network from another computer, or remote devices (such as a "media-player with remote control capabilities"). Some have theorized that perhaps the iPhone—or really any type of mobile phone—may be able to control the media center's menu capabilities via Bluetooth or some other wireless option.

Reading through the patent, there seems to be no limit to the number of modules that could be connected, opening up the doors for a truly versatile Apple TV solution with the simplicity of the current Front Row-like interface. The clear benefit to the segmented system with external modules—as opposed to an all-in-one device—is that it would allow customers to add on whatever extra functionality they prefer to the main device. This would allow power users to have all variety of extra modules for storing, playing, and streaming media—all through a centralized control hub—while more "average" users could settle for the simplicity of the main Apple TV-like device and just one or two extra modules as they see fit for their lifestyles.

It would seem to make the most sense to add these features onto the current Apple TV rather than launch a new device, as the Apple TV could make use of these technologies with just a few software updates and a hub for connecting the modules. The true question would be whether Apple would allow third party manufacturers to design modules for connection with the Apple TV, or whether the company would want to sell them with the Apple name themselves.

The patent describes a plugin system that would allow each of the modules to interface with the main multimedia hub in order to offer a consistent user experience on the receiving end. If Apple were to make the software available, third party manufacturers would likely fall all over themselves to make products to function with the Apple TV, just as they have for the iPod. This would save Apple the hassle of being tied into creating and maintaining the modules themselves. And with the potential to generate another thriving accessory ecosystem that revolves entirely around another Apple product, perhaps the Apple TV will end up becoming the "iPod for your TV" after all.