StreetReaderAI: Towards making street view accessible via context-aware multimodal AI

Admin

StreetReaderAI: Towards making street view accessible via context-aware multimodal AI

AI Chat builds on AI Describer and lets users ask questions about their current view, past views and nearby geography. The feature uses Google’s Multimodal Live API, which supports real-time interaction, function calling and temporarily retains memory of all interactions within a single session.

The system tracks and sends each pan or movement interaction along with the user’s current view and geographic context, including nearby places and current heading. Its context window is set to a maximum of 1,048,576 input tokens, which is roughly equivalent to over 4k input images.

Because AI Chat receives the user’s view and location with every virtual step, it can answer location-based questions using previous context from the same session. In the example given, a user can virtually walk past a bus stop, turn a corner and then ask, “Wait, where was that bus stop?” The agent can respond, “The bus stop is behind you, approximately 12 meters away.”

The feature is designed for situations where short-term session memory and spatial context matter, including navigation through a virtual environment or asking about something seen moments earlier.

Source: research.google.

Companies can share verified announcements through Newz9’s international press release submission page.