BLUF

As AR glasses, spatial computing devices, and multimodal AI systems become more practical, the next interface challenge is not only generating content but deciding how information should appear in physical space. I am developing a human-centered text-to-AR framework that turns menus, game instructions, tutorials, safety prompts, and service information into contextual, readable, low-friction spatial overlays.

Project Exploration. I am developing a human-centered interface framework for text-to-AR systems; the project focuses on interaction and information design, not a hardware-product guide.

Key takeaways

  • AR will not become useful only because glasses become smaller or models become smarter.
  • The real challenge is designing standards for how information should appear in physical space.
  • Text-to-AR conversion can transform menus, instructions, games, tutorials, and public information into contextual spatial overlays.
  • A good AR system should reduce cognitive load, not add another screen to the user’s face.
  • The future interface layer needs HCI rules around visibility, timing, attention, consent, and environmental context.

1) Why Text-to-AR Matters

Most digital information still arrives as text.

Menus are text.
Instructions are text.
Rules are text.
Game prompts are text.
Public notices are text.
Product labels are text.
Tutorials are text.

But the environments where we use that information are rarely text-first.

A restaurant is spatial.
A store is spatial.
A game is spatial.
A kitchen is spatial.
A gym is spatial.
A city is spatial.

This creates a mismatch.

The user has to move attention between the physical world and the screen. They read, interpret, translate, and then act.

Text-to-AR is the idea of removing that translation burden.

Instead of showing information as a flat page, the system converts it into a spatial overlay placed directly where the action happens.

A restaurant menu does not need to be a PDF.
It can become a live visual guide over the table.

A board game instruction manual does not need to be read separately.
It can become step-by-step guidance projected over the board.

A recipe does not need to sit on a phone screen.
It can become ingredient-aware prompts near the bowl, stove, and cutting board.

The opportunity is not only AR.

The opportunity is building a standard for how everyday text becomes usable spatial intelligence.


2) The Problem With Current AR Interfaces

Most AR demos still behave like floating phone screens.

They place windows in space, but the interaction logic remains old.

That is not enough.

A true AR interface should not ask:

Where should I place this screen?

It should ask:

What does the user need to know right now, and where should that information appear to support the action?

This distinction matters.

Bad AR increases mental load. It distracts, blocks vision, and makes the user manage another interface layer.

Good AR disappears into the task.

The user should feel:

I understood the environment faster.

Not:

I am using a headset.


3) Use Case 1: Restaurants

Restaurants are an ideal starting point because the information is structured but the experience is physical.

A text-to-AR system could convert a menu into:

  • ingredient overlays
  • allergy warnings
  • spice-level visualization
  • portion-size previews
  • dietary filters
  • price comparisons
  • pairing suggestions
  • cultural explanations
  • ordering guidance

Instead of reading a menu line like:

Paneer tikka with mint chutney and pickled onions

The user could see:

  • a small visual preview of the dish
  • the likely portion size
  • vegetarian confirmation
  • spice level
  • estimated protein/calorie range
  • ingredient substitutions
  • recommended pairing

The goal is not to overwhelm the user with data.

The goal is to make the decision easier.

A good AR restaurant layer should follow three rules:

Show only what matters

The system should not annotate everything.

It should respond to user intent.

A user with allergies needs a different overlay than a user optimizing for protein.

Preserve the social experience

Restaurants are social spaces.

AR should not pull the user away from the people at the table.

The interface should be glanceable, subtle, and optional.

Support confidence, not addiction

The system should help users make decisions faster, then disappear.

The best restaurant AR interaction might last 20 seconds.


4) Use Case 2: Games

Games are another strong environment for text-to-AR conversion because rules often prevent new players from enjoying the experience.

Board games, card games, tabletop games, and even physical sports drills often depend on written instructions.

An AR assistant could:

  • explain the next legal move
  • highlight relevant pieces
  • detect incorrect setup
  • show turn order
  • display strategy hints
  • visualize hidden mechanics
  • teach new players without stopping the game

For example, in a board game, the user could look at the board and ask:

What can I do next?

The AR layer could highlight three possible actions and explain each in one sentence.

This matters because the biggest friction in many games is not gameplay.

It is onboarding.

Text-to-AR can turn static rules into interactive coaching.


5) Use Case 3: Tutorials and Skill Transfer

A major opportunity for AR is turning instructions into embodied learning.

This applies to:

  • cooking
  • makeup
  • skincare
  • grooming
  • fitness
  • repairs
  • assembly
  • medical device usage
  • workplace training

Instead of watching a video and copying it, the user could receive real-time guidance over their own environment.

A skincare tutorial could say:

Apply this amount here.

A haircut tutorial could highlight:

Do not cut above this line.

A repair guide could show:

Remove this screw first.

This creates a new kind of human-computer interaction:

The interface does not only inform.

It guides action.


6) A Proposed Text-to-AR Pipeline

A practical system could follow a five-layer architecture.

1. Input understanding

The system ingests raw text from:

  • menus
  • manuals
  • PDFs
  • websites
  • signs
  • product labels
  • game rules
  • voice commands

The first job is to convert unstructured text into a structured task model.

2. Context detection

The system identifies the environment.

For example:

  • restaurant table
  • game board
  • kitchen counter
  • store shelf
  • bathroom mirror
  • street sign
  • classroom
  • workplace station

This can be done through computer vision, object recognition, location signals, and user intent.

3. Information prioritization

Not all text should become AR.

The system must decide what is relevant now.

A good prioritization engine should consider:

  • user goal
  • safety relevance
  • time sensitivity
  • physical location
  • task stage
  • visual clutter
  • confidence level

4. Spatial placement

The system chooses where overlays appear.

This is the hardest HCI layer.

Information should be close enough to the relevant object to make sense, but not so close that it blocks vision or action.

5. Interaction and correction

The user should be able to ask:

  • “Explain this.”
  • “Show less.”
  • “Translate this.”
  • “Is this vegetarian?”
  • “What do I do next?”
  • “Hide everything.”
  • “Why are you showing me this?”

The system should be controllable through voice, gaze, gesture, and minimal touch.


7) Toward an HCI Standard

The most important part of this idea is not the AR demo.

It is the standard.

If every AR application invents its own behavior, users will face interface chaos.

We need a shared design language for spatial information.

A possible standard could include:

Visibility rules

How much information can appear at once?

Distance rules

How far should overlays be from the user’s eye line?

Occlusion rules

What objects should never be covered?

Attention rules

When is the system allowed to interrupt?

What information can be detected, stored, or shared?

Confidence rules

How should uncertain model outputs be displayed?

Social rules

How should AR behave in shared spaces?

These rules matter because AR is not like a phone.

A phone interface is contained.

An AR interface enters the user’s perception of reality.

That makes HCI standards more important, not less.


8) Product Direction

A realistic MVP should avoid trying to build the full AR future.

The first version could be a mobile-based AR assistant.

MVP 1: Restaurant AR Menu Layer

User scans a menu.

The system produces:

  • dish previews
  • dietary tags
  • allergy warnings
  • ingredient explanations
  • personalized recommendations

MVP 2: Board Game AR Helper

User scans a game board or instruction sheet.

The system produces:

  • setup help
  • turn guidance
  • rule explanations
  • move validation

MVP 3: Tutorial Overlay Assistant

User uploads or opens a tutorial.

The system converts it into:

  • step-by-step visual guidance
  • object-aware prompts
  • progress tracking
  • safety warnings

These MVPs are narrow enough to test, but broad enough to prove the core interface pattern.


9) The Bigger Research Question

The deeper question is:

What happens when text stops being a document and becomes an environmental layer?

That shift changes how we think about search, learning, accessibility, and instruction.

Search becomes spatial.

Learning becomes embodied.

Accessibility becomes contextual.

Instructions become interactive.

The interface of the future may not be a chatbot, app, or website.

It may be an invisible translation layer between the information world and the physical world.

Text-to-AR is one step toward that layer.


10) Final Thought

The success of AR will not depend only on better hardware.

It will depend on whether the interface respects human attention.

A useful AR system should not decorate the world with data.

It should make the world easier to understand.

That is the real design challenge.

Not augmented reality for spectacle.

Augmented reality for comprehension.


References and further reading

  • Apple Human Interface Guidelines - Spatial Design: https://developer.apple.com/design/human-interface-guidelines/spatial-layout
  • Apple VisionOS Developer Documentation: https://developer.apple.com/visionos/
  • Meta Presence Platform: https://developers.meta.com/horizon/documentation/unity/unity-presence-platform-overview/
  • Microsoft Mixed Reality Design Guidance: https://learn.microsoft.com/en-us/windows/mixed-reality/design/
  • Google ARCore Documentation: https://developers.google.com/ar
  • W3C Immersive Web Working Group: https://www.w3.org/immersive-web/


  • GitHub profile: https://github.com/ChinmayA301
  • Portfolio / The Lab: https://app.chinmayarora.com
  • Project repository: https://github.com/ChinmayA301/text-to-ar-interface-standard
  • Prototype direction: mobile-first AR overlay assistant for menus, games, and tutorials
  • Future implementation stack: React Native, ARKit / ARCore, Vision-Language Models, OCR, spatial UI components

After the last line

Comments & sharing

Agree, disagree, add context, or send this to someone who would have a take.

Share

LinkedIn X Email

Public thread

Comments are public and attached to this post through GitHub Discussions. Sign in with GitHub to join the thread.

Open public discussions

Newsletter, eventually

Get the next brain dump.

No fixed cadence yet. Leave your email for the first issue when it exists.

One list. No schedule. Unsubscribe whenever.