Media generation tools
Four built-in tools that make media from a description: pictures, video, speech, and music. They need no API key and no configuration - assign one to an agent and ask. The agent chooses the settings for each request from what you asked for, so you steer them in plain language rather than in a form.
- Generate Image (
generate_image) - create an image from a description, or edit an existing one. - Generate Video (
generate_video) - create a short clip with sound, from a description or from a picture. - Generate Speech (
generate_speech) - turn words into spoken audio in a single voice. - Generate Music (
generate_music) - create a short instrumental piece.
To stitch these pieces into something publishable - a reel with captions and a voiceover, a text card, a carousel - pass the URLs to the Compose Media tool. To start from stock footage instead of generated media, see stock photo & video tools.
Setup
There isn't any. These are built-in tools: they take no credentials and have no configuration fields, so they appear under Built-in tools on the Tools page ready to assign. Open an agent, go to its Tools tab, and switch the ones you want on. Usage is billed to your workspace credit like any other model call.
Images
Ask for a picture and describe it. The more specific the description - subject, setting, lighting, style, mood - the closer the result. The image comes back in the reply, ready to view or download.
Editing an existing image
You do not need to start over to change something. Describe the change and the most recent image in the conversation is used as the starting point, whether you uploaded it or the agent made it. "Make the shirt red", "remove the background", "same scene at night" all edit that picture rather than inventing a new one. To edit a different image, give the agent its URL, or attach the file to your message.
Several versions at once
Ask for more than one version and you get them in a single reply, laid out side by side so you can compare them. Say which one you want to keep working on and the agent carries on from that image.
Video
Describe the scene, the motion, and the style, and you get a short clip. The clip has synchronised sound, so it is not silent.
Animating a picture
Point the agent at an image and describe how it should move, and the clip starts from that exact frame. It works on a picture the agent just generated ("now animate that") and on one you supplied. This is the reliable way to get a specific look on screen: settle the still image first, then animate it.
Speech and music
Generate Speech produces spoken audio in one voice. You can give it the exact words to read, or describe what you want said - "a warm twenty-second welcome message" - and let it write and read the line. You can ask for a particular voice. It is one speaker per clip; a two-person dialogue in a single clip is not supported, so generate each part separately and combine them.
Generate Music produces a short instrumental piece from a description of genre, mood, instruments, or tempo. It does not sing or speak - use Generate Speech for words.
Where the media goes
Everything generated is stored at a Hania media URL and shown in the reply: pictures inline, audio and video as players you can scrub. The URLs are ordinary links, so an agent can pass them straight to another tool - attach a picture to an email, post a clip to a social channel, or feed several into Compose Media.
Generated media is kept indefinitely. It is removed only when you delete the conversation it belongs to, which deletes its media with it.
Behaviors & gotchas
- Describe the picture, not the topic. "Person reviewing documents at a desk" beats "productivity article".
- Video takes minutes, images take seconds. Plan a reel around that: settle the stills first, animate once you are happy.
- Vertical output needs to be asked for. For a reel or story, say so - and remember that animating a picture inherits the picture's shape.
- Quality follows the description. These tools have no quality dial in the interface; detail in the request is what moves the result.
- Music is instrumental. For a voiceover over music, generate the speech and the music separately and combine them with Compose Media, which ducks the music under the voice for you.
Classification & lifecycle
All four are built-in tools that need no credentials, so there is nothing to rotate or revoke. They create new media rather than changing anything you already have, so the destructive and sends-data-externally flags do not apply and ship unset. Each is post-call-hook eligible. Note that a video started as a post-call action finishes after the call has ended, since generation runs in the background.