Multimodal models — systems that handle speech, text and images together — are what let Ray-Ban Meta glasses interpret the wearer's surroundings. A wearer can ask about the landmark in front of them, have text translated in place, and use a range of other features built on the same capability.

Shane, a research scientist at Meta, has worked on computer vision and multimodal AI for wearables for seven years. His team's research includes AnyMAL, a unified language model that reasons over text, audio, video and IMU motion sensor data. He joined Pascal Hartig on the Meta Tech Podcast to describe how foundational models for the glasses are built and what makes AI glasses a distinct engineering problem.

From acoustic signals to encoder zoos

AnyMAL's scope sets up one of the episode's recurring themes: the modalities a wearable can draw on go beyond speech, and acoustic inputs in particular carry more than words. Feeding such a range of signals into a single model means assembling many encoders, one per input type, and the conversation covers how that “encoder zoo” is organized.

Evaluation and iteration follow from there. The discussion takes in zero-shot performance, the cycle of iterating on models, and how LLM parameter size factors into the decisions a team makes.

What happens when a wearer asks a question

The middle of the episode turns to the request path itself: how a request originating from the glasses is processed end to end, and what changes when the input is moving imagery rather than a still. From there the constraints widen — serving billions of users, where the optimization headroom lies, and how user feedback is folded back into the models.

Open-source work and the Be My Eyes program also come up, along with what it's like to work alongside industry experts inside Meta.

Episode chapters

  • Intro 0:06
  • OSS News 0:56
  • Introduction Shane 1:30
  • The role of research scientist over time 3:03
  • What's Multi-Modal AI? 5:45
  • Applying Multi-Modal AI in Meta's products 7:21
  • Acoustic modalities beyond speech 9:17
  • AnyMAL 12:23
  • Encoder zoos 13:53
  • 0-shot performance 16:25
  • Iterating on models 17:28
  • LLM parameter size 19:29
  • How do we process a request from the glasses? 21:53
  • Processing moving images 23:44
  • Scaling to billions of users 26:01
  • Where lies the optimization potential? 28:12
  • Incorporating feedback 29:08
  • Open-source influence 31:30
  • Be My Eyes Program 33:57
  • Working with industry experts at Meta 36:18
  • Outro 38:55

Where to listen

The Meta Tech Podcast, produced by Meta, focuses on the work its engineers do across every level, from low-level frameworks up to end-user features. The episode is available on Spotify, Apple Podcasts, Pocket Casts, Overcast, or wherever else podcasts are found. Feedback goes to the show on Instagram, Threads or X.

Links