“I see what you mean”: how screen-aware avatars transform digital journeys
Why we decided to give our Agentic Avatars eyes.
Agentic Avatars were designed to make enterprise knowledge easier to access and act on. The idea was simple: what if people could interact with an organization’s accumulated content and data as naturally as they interact with another person?
Instead of searching through a vast knowledge base, navigating a course catalog, or finding the right document, users can simply ask an agent that knows the organization’s content, products, policies, processes, and expertise.
But as we thought more about how people actually use conversational agents, it became clear that something was missing. Sometimes an agent needs to understand not only what a user is asking, but what is happening on the user’s screen.
Consider a learner who gets stuck halfway through a new process, a marketer doing a final check before sending a high-stakes email campaign, or a customer configuring a product who can’t work out why the next step is unavailable. In each case, the organization may already have the knowledge needed to help. What’s missing is the current state of the task.
A human expert sitting beside you has a simple advantage in these situations: they can look over your shoulder. We wanted our Agentic Avatars to have that advantage too. So we decided to give them eyes.
Why seeing matters when it comes to workflows
There is a big difference between an agent that answers questions and one that helps people get work done.
If a marketer asks, “What is our policy for sharing customer information externally?”, a knowledge base may be enough to provide a useful answer. But if they are halfway through reviewing a document and ask, “Is this safe to share?”, the question can’t really be separated from what is on the screen.
The same is true in many workplace and customer experiences. A learner may know the process but miss one incorrect field. A marketer may understand the campaign brief but accidentally select the wrong audience. A customer may follow setup instructions correctly until one configuration choice puts them on a different path.
What’s on the screen often tells you where someone is in a process, what they have already done, which options are available, and where they may be stuck. Without that visual context, an agent has to reconstruct the situation indirectly by asking the user to describe the page, explain previous choices, or upload screenshots, adding friction and ambiguity.
Visual context lets the user show the agent the situation instead of translating it into words.
What screen sharing adds to an Agentic Avatar
Screen sharing adds a vital layer to human-agent interactions. With the user’s permission, the Agentic Avatar can see the screen, application, or browser tab the user chooses to share and use that visual information as part of the conversation.
The agent can then bring together three things:
Enterprise knowledge: what the organization has taught it about its products, processes, policies, content, and expertise.
Conversation: what the user is asking, saying, and trying to accomplish.
Visual context: what the user is looking at in that moment.
The screen can show the agent that a marketer has selected “All Customers.” The organization’s knowledge can tell it that the campaign brief calls for prospects in a particular region, while the conversation provides the intent behind the final check.
Seeing without relevant knowledge is observation. Enterprise knowledge without an understanding of the immediate situation is general advice. Bringing the two together makes it possible to provide guidance tied to the task in front of the user.
From answering questions to understanding progress
Screen awareness also changes the kinds of questions an agent can meaningfully help with.
Take a new employee learning a complex internal application. They can ask an Agentic Avatar, “How do I create a new request?” and receive in-the-moment instructions based on the organization’s training material.
Now imagine they are halfway through the workflow. They have completed several steps and reached a screen that doesn’t match the training example. Their question changes from “How does this work?” to “What am I missing?”
The first is primarily a knowledge-retrieval problem. The second requires an understanding of what the learner has already done and what is currently in front of them.
With the screen in view, the agent can ground its guidance in that current state. Perhaps the employee overlooked a mandatory field or selected an option that changed the next step. Instead of repeating generic instructions, the conversation can focus on the point where help is actually needed.
Real workflows are rarely as linear as training materials make them look. People make different choices, work in changing interfaces, and get stuck in ways documentation cannot always anticipate.
Learning & Enablement: closing the gap between knowing and doing
For Learning & Enablement teams, this opens up possibilities beyond the traditional course-and-assessment model.
One of the hardest problems in workplace learning is transfer: whether people can apply what they learned when the real situation arrives. A learner may understand a policy during training and still overlook the relevant detail when applying it weeks later.
Imagine a privacy and compliance exercise in which an employee has to decide whether a document is ready to share with an external partner. They understand the policy and have removed the obvious personal information. Before submitting their answer, they share the screen and ask the avatar to review their work.
The avatar can see that customer information is still visible inside an embedded screenshot and connect that detail to the organization’s privacy guidance. The value is not simply that AI spotted some text; it is that the agent can explain why that information matters in this organization and in this exercise.
The same model can apply to software training, process training, onboarding, certification, and other forms of practice. Learners can interact with a personal AI guide throughout the process as avatar provides instructions or advice informed by both the organization’s knowledge and what is happening on screen.
For Learning leaders, that moves the conversation beyond course completion toward helping employees apply knowledge in the moments that matter.
Revenue Engagement: helping teams apply what they already know
Revenue organizations face a similar problem from a different direction.
Sales and marketing teams rarely suffer from a shortage of information. They have messaging frameworks, campaign briefs, product updates, competitive intelligence, pricing guidance, sales playbooks, customer research, and brand guidelines. The difficult part is applying the right piece of that knowledge at the point of execution.
Imagine a marketer preparing an email campaign on an emailing platform. The copy has been approved, the creative is ready, and they are reviewing the setup one last time before sending it to 48,216 recipients. They ask the Agentic Avatar, “Does this look right?”
The avatar can see that “All Customers” is selected as the audience. Because it also has access to the campaign knowledge, it understands that the email was created for prospects rather than existing customers and can flag the discrepancy before it is sent.
That is different from giving the marketer another place to search for the campaign brief. The problem is not that the correct information doesn’t exist. It is connecting that information to the decision being made on screen.
The same principle can apply to a salesperson reviewing a customer-facing presentation, a rep practicing a product workflow, or an enablement user choosing content for a particular situation. For us, that is an important direction for Revenue Engagement: moving from distributing knowledge to helping teams apply it.
What other use cases benefit from this capability?
The same product logic extends to customer and partner experiences.
A customer onboarding into a complex service may reach a configuration screen where they do not know which setting is causing the problem. Instead of trying to describe the interface, they can show it. The Agentic Avatar can then guide them to the solution and suggest further settings variables to prevent them getting stuck again. A partner configuring an integration may already have access to excellent product documentation, but visual context can help the agent understand where they are in the process and which guidance is relevant.
In both cases, the organization already knows how to help. The challenge is delivering the right knowledge with enough context to make it useful.
Building agents that understand more of the journey
Enterprise agents are becoming more capable, but the real test is whether they can be useful inside the journeys where employees, customers, and partners actually need them.
Knowledge is one part of that. Conversation is another. Visual context adds an understanding of the task in front of the user.
When those pieces work together, users spend less time translating their problem into a format the AI can understand. They can show the agent what they mean and continue the conversation from there.
That is what makes screen sharing interesting to us. It is not sight for sight’s sake. It is a way to make the intelligence already behind the Agentic Avatar more relevant to the workflow in front of it.
Sometimes the difference between knowing the answer and being genuinely helpful is simply being able to see the same thing.
See Kaltura Agentic Avatars in action
Discover how Kaltura Agentic Avatars bring enterprise knowledge, natural conversation, and real-time visual context together to provide guidance in the moments that matter.
Frequently Asked Questions
What is screen-aware AI?
Screen-aware AI can use visual information from a screen that a user chooses to share as additional context for an AI interaction.
This allows the AI to consider not only what the user says, but also relevant information visible in the digital experience they’re discussing.
Why did Kaltura add screen sharing to Agentic Avatars?
Agentic Avatars already combine conversational interaction with enterprise-specific knowledge. Screen sharing adds another important source of context: what the user is seeing at that moment.
The goal is to make the knowledge available to the avatar more actionable by connecting it to the user’s immediate situation.
What is a screen-aware Agentic Avatar?
A screen-aware Agentic Avatar combines enterprise knowledge, conversational AI, an interactive avatar experience, and visual context from a user-shared screen.
Together, these capabilities can enable more contextual guidance than conversation or screen sharing alone.
How is this different from a traditional chatbot?
A chatbot generally relies on the information contained in the conversation and its available knowledge sources.
With screen sharing, the Agentic Avatar can also use information visible on the user’s shared screen as context. This can reduce the need for the user to describe where they are, what they’re seeing, or what they’ve already done.
Why is enterprise knowledge important for screen-aware AI?
Seeing something and understanding its significance are different things.
Enterprise knowledge allows the Agentic Avatar to interpret visual context using information specific to the organization, such as policies, training materials, processes, product knowledge, messaging, and approved content.
How can screen-aware Agentic Avatars support Learning & Enablement?
Screen awareness can help connect formal learning with application.
A learner can share the task or simulation they’re working on and receive guidance informed both by what’s on screen and by the organization’s learning content and knowledge. The Agentic Avatar can function as a personal tutor, ensuring knowledge retention and correct application.
This can support more contextual feedback and learning closer to the moment of need.
How can screen-aware Agentic Avatars support Revenue Engagement?
Revenue teams need to apply constantly changing product, messaging, campaign, and customer knowledge.
Visual context can help an Agentic Avatar connect that knowledge to the specific workflow, content, or decision a user is working on at that moment.
Can an Agentic Avatar see everything on a user’s computer?
No. Screen sharing is user initiated. The user chooses the screen, application window, or browser tab they want to share and can stop sharing.
Does the Agentic Avatar take control of the user’s screen?
Screen sharing is designed to provide visual context for the conversation. The user remains in control of their own workflow and actions.
Why does visual context matter?
Often, the information required to answer a question is already visible, but difficult for the user to describe or easy to overlook.
Visual context can help the Agentic Avatar understand where the user is, what they’re doing, and which part of its enterprise knowledge is relevant to the situation. In addition, the user may bring information to the conversation that the agent does not have, and by sharing their screen, in effect teach the Agentic Avatar new knowledge.
Was this post useful?
Thank you for your feedback!