This repository demonstrates how to build a video generation agent using the Google Agent Development Kit (ADK), Gemini 2.5 Flash Image (Nano Banana), and Veo 3.1.
It is a full-stack web application designed to be deployed on Google Cloud Run, with ADK Web UI, using Vertex AI Agent Engine Sessions Service for session management and Google Cloud Storage for storing artifacts.
- Story Generation: Creates a story with a plot and character descriptions.
- Storyboard Creation: Generates storyboards for each shot of the story, including first and last frames.
- Video Generation: Produces video clips for each shot using the generated storyboards.
Capybara dancing with a banana 🍌 - a keyframe image for Veo 3.1, generated with Nano Banana
The video generation agent is composed of a main agent and three sub-agents:
- Root Agent (Orchestrator): The main agent that orchestrates the video generation process. It takes user input and delegates tasks to the appropriate sub-agent.
- Story Agent: Responsible for creating the story, including the plot and character descriptions.
- Storyboard Agent: Generates the storyboard for each shot, including the first and last frames. It uses the Gemini 2.5 Flash Image model (Nano Banana) to generate the images while preserving character and scene consistency.
- Video Agent: Creates the video for each shot using the Veo 3.1 model. It takes the first and last frames from the storyboard and generates a video that transitions between them.
The Orchestrator performs delegation to sub-agents through various stages - from building the story to generating videos. The reason for using LLM-driven delegation rather than a SequentialAgent for a sequential workflow is because the process is iterative, and the user may potentially jump multiple steps back to make corrections in the story or media generation prompts.
- An existing Google Cloud Project. New customers get $300 in free credits to run, test, and deploy workloads.
- Google Cloud SDK.
- Python 3.11+.
-
Clone the repository:
git clone https://github.com/vladkol/media-generation-agent.git cd media-generation-agent -
Create a Python virtual environment and activate it:
We recommend using
uvuv venv .venv source .venv/bin/activate -
Install the Python dependencies:
uv pip install pip uv pip install -r agent/requirements.txt
-
Create a
.envfile in the root of the project by copying the.env-templatefile:cp .env-template .env
-
Update the
.envfile with your Google Cloud project ID, location, and the name of your GCS bucket for AI assets.
To run the agent locally, use the run_local.sh script:
./deployment/run_local.shThis will:
- Register an Agent Engine resource for using with the session service.
- Start a local a web server with the ADK Web UI, which you can access in your browser.
To deploy the agent to Cloud Run, use the deploy.sh script:
./deployment/deploy.shThis script will:
- Register an Agent Engine resource for using with the session service.
- Deploy the agent to Cloud Run, with the ADK Web UI.
The agent's behavior is defined by the prompts in the agent/video_generation/prompts directory.
story_agent.md: This prompt instructs the agent on how to create a story, including the plot and character descriptions.storyboard_agent.md: This prompt guides the agent in creating a storyboard for each shot, including generating the first and last frames.video_agent.md: This prompt tells the agent how to generate a video for each shot using the storyboard frames.
The agent uses the following tools:
nano_banana_tool.py: A tool for generating images using the Gemini 2.5 Flash Image model.veo3_agent.py: A tool for generating videos using the Veo 3.1 model.