InAct Entertainment

Real-time music models: when the listener becomes part of the instrument

Explore the musical feedback loop: intention, control, audible response, and a better next decision.

Back to courses
Start here · Plain-language foundations

Before you begin

An ordinary player plays a recording that already exists. A real-time music tool can create new sound while you use its controls. You might ask for more rhythm or a different texture, then listen to what happens. The important question is whether the change matches what you meant. This lesson studies the interaction, not just how impressive the music sounds.

Words you’ll use in this lesson
Generative
creating new material from learned patterns.
Texture
how the sounds combine and feel together.
Density
how much musical activity is present.
Feedback loop
act, notice the result, and decide what to do next.
A useful result

What you’ll be able to do.

ExplainApplyCheckBuild
Try a decision · Illustrative example

What would you change first?

A music app has many controls, but a new listener cannot tell which change affects the sound.

Pause and choose your first action. Explain why it would help before opening the example response.

Compare your reasoning →

Test one control at a time. Ask the listener to describe what changed and whether it helped their musical goal. Use that evidence to improve the control labels.

Your learning route

From choosing a song to shaping a musical moment

Imagine opening a creative workspace and hearing a quiet pulse. You turn one control, and the texture becomes warmer. You turn another, and percussion gradually enters. Instead of searching a catalog for the perfect track, you participate in how the sound develops. This imagined experience illustrates a different relationship with generated music: the listener becomes a director of an unfolding moment. The attraction is not simply that the music is new. The attraction is that a choice can become something you hear, evaluate, and change again.

Real-time music models make that relationship worth studying. Google DeepMind describes Lyria RealTime as an interactive system that produces a continuing musical stream and responds to controls for musical attributes. Its public model page describes prompt blending and controls including tempo, density, brightness, and key. These are advertised capabilities of a particular system, not a promise that every music generator or account offers identical controls. (Google DeepMind, n.d.)

The research paper Live Music Models introduces continuous generation with synchronized user control and presents Magenta RealTime alongside the API-based Lyria RealTime. The authors describe an emphasis on human participation during live creation. The paper was first submitted in August 2025 and revised in November 2025. Its research framing is useful context, rather than evidence that every future application will feel responsive, enjoyable, or commercially successful. (Lyria Team, 2025)

This lesson develops an original design method around that possibility. We will separate a musical idea from a control, separate technical responsiveness from a listener’s experience, and turn an exciting demonstration into a small experiment that can be evaluated. The examples are hypothetical. You do not need an API account, programming background, or paid music product to complete the planning exercises. A paper prototype and an ordinary recorded track can teach much of the design thinking before anyone connects a generative system.

What the listener is actually controlling

A playback interface controls a recording. Play, pause, seek, and volume change how an existing file reaches the listener. A generative interface may change what the system creates next. The controls can look similar while doing different jobs. A volume slider makes an existing sound louder or softer. A density control could influence the amount of musical activity. Confusing those jobs makes the interface feel unpredictable, especially when a listener expects an immediate physical response from a familiar-looking slider.

Start by writing the promise beside each proposed control. A control named Energy might mean additional percussion, faster musical motion, a brighter timbre, or merely greater loudness. Those possibilities are related in everyday language but distinct in practice. If the design cannot describe what a control is intended to influence, it will struggle to explain why the result changed. Prefer a narrow promise at the beginning. A clearly explained percussion control can be more engaging than an ambitious Mood control that appears to do everything.

IntentSoundDecisionNOTICE · EVALUATE · ADJUST
The listener acts, hears, evaluates, and shapes the next moment.

Next, identify the boundary between intent and result. A user can request more intensity without receiving an exact predetermined arrangement. The system may interpret that request in a way the designer did not anticipate. Your interface should communicate direction without suggesting complete command over every musical event. A helpful label might explain that changes shape the upcoming texture. It should not imply that a slider guarantees a specific chord or instantly rewrites audio that has already been heard.

Practice with an imaginary two-control instrument. One control adjusts the presence of rhythmic activity. The other blends a soft, spacious character with a sharper, more percussive character. Describe what happens at both ends of each control. Then describe one thing that must remain stable, such as a comfortable listening level or an overall musical identity. That stable element is the anchor. Without an anchor, every movement may feel like switching to an unrelated song rather than participating in the same evolving piece.

A musical control needs an audible relationship with the listener’s intention.InAct editorial insight
Learning checkpoint 1 · Choose and explain

Connect the ideas

Try each question. The feedback explains the idea so you can apply it.

What distinguishes a generative control from volume?
What is a useful first instrument?
What follows an audible response in the learning loop?

0 / 3 correct

Design the smallest useful musical loop

A useful loop has four parts: an intention, an action, an audible response, and a decision. The listener wants a quieter texture, moves a control, hears the response, and decides whether it fits. If any part is unclear, adding more controls usually increases confusion. A beautiful panel is not enough. The listener must understand what to try, notice something relevant, and retain the ability to adjust or stop. This is a design problem as much as a generation problem.

In a first prototype, give the listener a short task. Ask them to create a background for reading that feels spacious but does not become distracting. That task provides a listening context and a reason to evaluate. It also makes success observable. A participant can explain whether the result supported reading and which changes helped. Asking whether the instrument was cool may produce enthusiastic reactions, but it tells you little about whether the controls matched the person’s intention.

Resist the temptation to make the first experiment a complete performance platform. You do not need a genre browser, a social feed, recording tools, audience voting, and an elaborate avatar to learn whether two controls make sense. Each additional feature creates another possible explanation for confusion or enjoyment. Choose one context and one small musical relationship. Learn how people interpret that relationship, then expand the instrument with evidence instead of assuming that more interaction creates more value.

Document the loop in ordinary language. For example: the listener wants less rhythmic activity, lowers the percussion control, waits for the upcoming texture, and judges whether the pulse is calmer. Now add the interface response. The control should visibly acknowledge the action, even if the musical change arrives later. A status message can say that the next texture is being shaped. It should not pretend that generation is finished simply because the interface has accepted the request.

Make timing part of the experience

Responsiveness is not one number. There is the delay before an interface acknowledges a movement, the delay before a system accepts a new instruction, and the delay before the listener notices a musical change. Those moments can differ. A control might respond visually right away while the audible result emerges over several beats. The design should help the listener understand that relationship. Otherwise, they may move the control repeatedly, assuming the first instruction was lost, and create a sequence of changes they never intended.

A practical experiment records both the action and the perceived response. Ask a participant to mark when they moved the control and when they first heard a relevant change. Do not present an invented universal latency threshold as a fact about music systems. Establish a target appropriate to the task and equipment, then test it. A performer trying to create an abrupt gesture may need a different experience from a reader shaping a slowly changing background. The context determines what kind of timing matters.

Use gentle transitions deliberately. A gradual adjustment can sound intentional when the interface explains that the upcoming music is being shaped. The same adjustment can feel broken if the interface promises an instant switch. Think about the relationship between the name of the action and its audible pace. Words such as blend, drift, or shape can communicate gradual influence. Words such as cut, hit, or switch suggest a more immediate event. Labels are part of the instrument’s musical contract with its listener.

Plan what happens when requests arrive faster than useful musical responses. The system might keep the newest intent, limit the rate of changes, or allow a deliberate reset. The right choice depends on the actual model and application. At the design stage, make the policy explicit rather than hiding it behind a responsive animation. A moving control is only visual feedback. It is not evidence that every instruction has become a distinct audible change. Your prototype should teach that difference instead of disguising it.

Watch · Notice · Apply

Watch an optional perspective.

The written lesson stands on its own. These optional videos offer another example or perspective. Read the focus, watch if you choose, and try the idea yourself. If a video is unavailable, continue with the explanation and practice on this page.

Creative demonstration

Jacob Collier x Gen Music | Google Lab Sessions

Jacob Collier · Google

Start at 00:00. Notice the human choices while sound changes. This is an experimental tool demonstration, not a release-rights guide.

Watch on YouTube at 00:00 ↗

Google. (n.d.). Jacob Collier x Gen Music | Google Lab Sessions [Video]. YouTube. Original publisher link above. Video and availability remain controlled by the publisher.

Videos load only when you choose to watch. The original publisher controls playback; use the direct link if a player is unavailable. Captions and playback speed are available when the publisher provides them.

A hypothetical listening-room experiment

Consider a small team building an interactive listening room for a portfolio website. The aim is to let visitors shape a background texture while reading about the studio’s work. This is a hypothetical project, not a description of a service currently available on the InAct business card. The team chooses two controls, Space and Pulse, plus clearly visible play, pause, and volume. It also decides that listening must begin through an intentional user action rather than surprising someone with sound when the page opens.

Before generating any audio, the team builds a paper version. Participants point to the controls and explain what they expect. One person assumes Space changes stereo width. Another assumes it adds silence. A third expects a larger room effect. Those interpretations reveal a naming problem that a working model would not automatically solve. The team replaces the vague name with Texture, adds a short explanation, and tests again. This inexpensive step prevents technical work from being mistaken for communication work.

IntentSoundDecisionNOTICE · EVALUATE · ADJUST
The listener acts, hears, evaluates, and shapes the next moment.

The team then uses a small set of recorded examples to simulate possible responses. A facilitator changes the recording after a participant moves a control. This simulation does not prove a model can reproduce the same transition, but it helps test the intended interaction. Participants describe whether they hear more rhythmic activity and whether the progression feels coherent. The facilitator records misunderstandings, not merely positive reactions. A useful experiment can reveal that an attractive feature should be simplified or postponed.

Only after those steps does the team connect a real generative system. The technical test asks whether the available controls can support the interface promise with acceptable continuity. The design test asks whether visitors can use the experience while still reading. Those tests are related but separate. A model may generate impressive music while the website becomes distracting. Conversely, a calm interface may feel intuitive while the generated transition fails to communicate the requested change. Evaluate both before calling the experience successful.

Learning checkpoint 2 · True or false

Test the assumptions

Try each question. The feedback explains the idea so you can apply it.

A moving slider proves every request became audible.
A paper prototype proves a music model can perform the transition.
Silence is a legitimate listener choice.

0 / 3 correct

Measure the relationship between choice and sound

Create a compact observation sheet with four questions. Did the participant understand the control? Did they hear a relevant change? Could they explain the relationship between their action and the sound? Could they recover when the result was not what they wanted? These questions turn a vague impression into specific observations. You can still ask about enjoyment, but enjoyment should sit beside evidence about control and understanding. It should not replace those measures.

In an illustrative test with ten participants, suppose eight can explain the Pulse control before using it, six notice its intended effect, and four can reliably return to a calmer texture. These invented counts are teaching data, not measured client results or a benchmark for real-time models. The drop across stages suggests a recovery problem. The team should inspect how people return to a previous state rather than conclude that the interface is ready because most participants understood its label.

Listen for conflicting interpretations in participant explanations. Someone may correctly hear additional percussion yet describe it as louder because the whole arrangement feels more intense. That observation does not mean the person is wrong. It means your explanatory language should connect technical control with ordinary listening language. Ask what changed, what stayed stable, and what the person wanted next. These questions help distinguish an unintended result from a result that was useful but described differently from the designer’s vocabulary.

Keep a decision log after each session. Write the observation, your interpretation, the proposed adjustment, and the next test. An observation is that three people repeatedly moved the same slider before hearing a change. An interpretation is that they expected a faster response. The adjustment might be a clearer status message or a different musical transition policy. The next test checks whether repeated movements decrease. Keeping those elements separate prevents an attractive explanation from becoming an unsupported claim of improvement.

Interactive planning exercise

Give your judgment a shape.

Adjust the sliders to compare your own assessment. These starting values are illustrative, not measured results or industry benchmarks. Explain your scores in the application notes.

Give the listener a way back

Experimentation feels safer when mistakes are reversible. A listener should be able to pause, reduce volume, reset the current direction, or return to an earlier starting point when the system supports that behavior. Do not imply exact audio recovery if the model cannot reproduce a previous stream. A reset can restore control settings without restoring the precise sound already generated. Explain what is being restored. A small distinction in wording can prevent a large misunderstanding about how the musical instrument works.

Treat silence as a legitimate choice. Some visitors want an animated site but do not want audio. Others will use assistive technology, work in a shared room, or read on a device without headphones. The music experience should be available without dominating the page. A clear stop control and stable navigation support the main purpose of the website. If the sound is meant to accompany reading, the reader’s ability to continue reading matters more than maximizing time spent inside the player.

Provide keyboard access and visible focus for every interactive control. Give controls names that make sense without relying on their shape or color. A decorative equalizer can indicate activity visually, but a person should also have access to meaningful playback status. Avoid communicating success, loading, or errors through animation alone. A readable message and a controllable interface make the experience more robust. These design requirements can be planned before the underlying music generation is connected.

Consider a quiet mode for the visual experience as well. Constant motion in a player, a graph, and a website background can compete for attention. A reduced-motion preference should simplify decorative animation without removing essential information. This does not require making the instrument dull. A carefully chosen movement can communicate that sound is active, while the rest of the interface stays calm. Distinguish functional feedback from decorative activity and give each a deliberate purpose.

Keep records without confusing them with rights

A useful session record can include the model and version when available, the date, the control settings, the prompts, the user’s decisions, and the intended use. Such a record helps a team reconstruct its creative process and investigate differences between sessions. It does not automatically establish ownership, permission to distribute, or an entitlement to monetize the result. Those questions depend on applicable terms, the materials involved, and the circumstances of the work. A technical record and a rights determination are different things.

For U.S. copyright context, the Copyright Office’s January 2025 announcement explains that protection for AI-related output depends on sufficient human expression and that prompting alone does not establish the required authorship. That is a limited summary of a specific U.S. source, not a legal conclusion about your session or every jurisdiction. (U.S. Copyright Office, 2025) If you plan a commercial release, examine the current product terms and the relevant distribution requirements separately from the artistic question of whether the music sounds compelling.

IntentSoundDecisionNOTICE · EVALUATE · ADJUST
The listener acts, hears, evaluates, and shapes the next moment.

Keep the record readable. A folder full of unnamed exports and copied prompts is difficult to use when a collaborator asks which version was approved. Name sessions by purpose and date, record the intended audience, and identify the person responsible for reviewing the final result. If several people contribute to a live interaction, describe their roles accurately. Avoid transforming a technical log into a dramatic story about a human performance that did not take place. The experiment can be interesting without invented provenance.

For a prototype that uses recorded examples instead of live generation, say so in the testing notes. That distinction protects the quality of your evidence. You may have learned that listeners understand a control, while still needing to verify whether a model can support it. Honest records make those next steps obvious. A polished demo should not erase the difference between simulated interaction, tested integration, and a production experience that has been observed over time.

Learning checkpoint 3 · Scenario decisions

Choose the next move

Try each question. The feedback explains the idea so you can apply it.

People move a slider repeatedly before hearing change. What next?
Reset restores settings but not identical audio. What should the label do?
Most people understand a label but cannot recover. What next?

0 / 3 correct

Build a useful first-session lesson

Teach the instrument through a small sequence rather than a wall of instructions. Begin with one stable texture and invite a single change. Ask the listener to notice the result. Then invite them to return toward the starting direction. Finally, let them combine two controls. This progression gives the person an opportunity to develop a relationship with the instrument before navigating a complicated menu. The lesson is embedded in participation, not saved for a help page that few visitors will open.

Use a listening prompt that encourages attention. Instead of asking whether the result is better, ask where the rhythmic activity became more noticeable or whether the transition supported the intended mood. Better has no stable meaning without a purpose. A sound that is useful for a dramatic entrance may be poor for sustained reading. Your lesson should show that creative evaluation depends on context. The same observation can guide a musician, a designer, or a person shaping a personal listening environment.

Create a simple comparison exercise using one controlled difference. Hold the musical intention steady while changing the amount of pulse. Ask the listener to identify which version supports a chosen activity and explain why. There is no need to treat the exercise as a universal test of taste. The goal is to practice connecting intention, audible evidence, and a decision. When participants can explain that relationship, they have learned something transferable beyond the particular tool.

End the session by asking the listener to describe an experiment they would try next. They might blend two textures, explore a quieter rhythmic direction, or compare how the same controls feel in a different context. Their answer can reveal whether the interface taught a usable mental model. Someone who only says they would press random buttons may have enjoyed the novelty without understanding the relationship. That is useful feedback for the next design iteration, not a reason to judge the participant.

Watch · Notice · Apply

Watch an optional perspective.

The written lesson stands on its own. These optional videos offer another example or perspective. Read the focus, watch if you choose, and try the idea yourself. If a video is unavailable, continue with the explanation and practice on this page.

Technical overview

Opportunities in AI, 2023

Andrew Ng · Stanford Online

Start at 00:00. Use the talk as a foundation. Write down one task, one possible use of AI, and one result you would need to check. Product examples are from 2023.

Watch on YouTube at 00:00 ↗

Stanford Online. (n.d.). Opportunities in AI, 2023 [Video]. YouTube. Original publisher link above. Video and availability remain controlled by the publisher.

Videos load only when you choose to watch. The original publisher controls playback; use the direct link if a player is unavailable. Captions and playback speed are available when the publisher provides them.

Decide what is ready and what remains experimental

A responsible launch decision uses specific evidence. Confirm that the core controls communicate their purpose, the audible response supports that purpose often enough for the intended experience, and visitors can stop or recover without confusion. Check behavior when the service is unavailable, a session disconnects, or the device cannot play audio. Write down the fallback. A silent page with a clear explanation is better than an interface that appears active while giving the visitor no useful result.

Review the business purpose alongside the music. Does the experience help visitors understand the studio, explore its portfolio, or discover an original experiment? If it only distracts from the work, more elaborate generation may not solve the problem. Choose measures that connect to the intended purpose, such as whether participants can find a destination after listening or describe the studio’s creative approach. Avoid treating session length as proof of value. Long sessions can reflect curiosity, confusion, or difficulty leaving.

Build an experimental release in stages. Start with an internal prototype, then a small invited test, then a clearly described public experience if the evidence supports it. Record what is simulated, what is connected, and what is dependable. Update the description as those boundaries change. Real-time music is an opportunity to explore a new kind of participation, but participation becomes meaningful through the relationship between a person’s intention and the sound that follows.

The deeper lesson is that an evolving soundtrack is also an evolving conversation. The designer chooses what can be expressed through the interface. The listener chooses what to try. The system contributes musical material, and the listener decides what that material means in context. When those roles are clear, the experiment can teach more than how to operate a generator. It can teach how to listen with intention, how to evaluate a creative system, and how to shape an experience without pretending every surprise is under complete control.

Apply it to your work

Your next experiment.

Sketch a two-control music experience. Define the intended change, the audible evidence, a recovery action, and one test with a listener. Identify what is simulated and what would require a working model.

Final assessment

Bring it together.

Apply the ideas, then review your assessment. You can retry any answer.

What makes a useful launch measure?
What does a session log NOT automatically establish?
Which prototype sequence creates useful evidence?

References

Google DeepMind. (n.d.). Lyria RealTime. https://deepmind.google/models/lyria/lyria-realtime/

Lyria Team. (2025). Live music models [Preprint]. https://doi.org/10.48550/arXiv.2508.04651

U.S. Copyright Office. (2025, January 29). Copyright Office releases Part 2 of Artificial Intelligence Report. https://www.copyright.gov/newsnet/2025/1060.html

Sources checked October 9, 2026. Linked author-date citations identify reported developments. Frameworks, exercises, and labeled hypothetical examples are original InAct editorial guidance. Undated pages use n.d.; platform availability and policies can change.

Make the lesson yours

Your application notes.

Notes and article bookmarks stay in this browser on this device. No account is needed. They are not sent to InAct. Clearing browser data can erase them.