Google DeepMind rethinks the mouse cursor in the era of generative AI
Google DeepMind researchers Adrien Baranes and Rob Marchant design a smart cursor prototype powered by Gemini to bring generative AI directly to any screen.
Google DeepMind introduces a smart cursor prototype powered by Gemini, designed to extend the pointer's functionality beyond a simple on-screen coordinate. The idea championed by researchers Adrien Baranes, PhD. and Rob Marchant: to bring AI to where the user works, rather than forcing them to switch to a dedicated application for each request.
The system is based on four interaction principles. The first is to maintain the workflow: the cursor can be invoked in any application to summarize a PDF into bullet points, transform a table into a pie chart, or double the quantity of ingredients in a recipe. The second relies on the automatic capture of visual and semantic context around the pointer, which eliminates the need to write a detailed prompt. The third leverages everyday language, consisting of "this" and "that," combined with gestures. The fourth transforms pixels into actionable entities: a photographed handwritten note becomes an interactive to-do list, and a still image from a travel video leads to a restaurant reservation link.
The initial practical applications are integrated into Chrome, where Gemini can now be prompted on a specific portion of a web page (comparing multiple selected products, projecting a sofa into one's living room). A feature dubbed Magic Pointer is also set to arrive on Googlebook, the company's new laptop, and further tests are planned in Google Labs Disco. The prototype is currently available for use in Google AI Studio for two use cases: image editing and searching for locations on a map.