02Deploy AI systems

Lecture Video Generation Application

The model worked. The infrastructure was the problem.

An AI video generation application designed to run GPU-heavy workloads without keeping expensive infrastructure idle.

0
idle GPU cost
5
lifecycle boundaries
Hours
lecture production, not days

Focus

Deployment · Cloud · GPU · Serverless · Cost · Operations

Stack

AWS · RunPod Serverless · Docker · VietTTS · MimicTalk · Stable Diffusion

Context

The AI pipeline could already generate the individual pieces of a lecture video.

  • Text could become speech.
  • A teacher image could become a talking video.
  • Slides could become visual content.

The next problem was turning these capabilities into an application that could actually be deployed and operated.

The workload was GPU-intensive and long-running. A single video could require several inference stages, while GPU resources would sit idle whenever nobody was generating a video.

So the deployment question became:

How do I build an infrastructure model that matches the actual workload instead of treating GPU inference like a normal API?

Architecture

The deployment architecture

I separated the always-available application layer from the compute-intensive AI workloads.

  1. Web / API Layer

    AWS · always available

    always on
    submit job
  2. Job Processing

    request lifecycle ends here

    always on
    on-demand GPU
  3. RunPod Serverless

    VietTTS → MimicTalk → …

    on demand
    write artifact
  4. Storage

    artifacts outlive workers

    persistent
    serve
  5. Output

    persistent

The important part wasn't any individual model.

It was deciding where each part of the system should run and when it should consume compute resources.

Engineering decisions

The engineering decisions

01

Separate the application layer from GPU inference

Problem

The application needed to remain available even when no video was being generated.

However, the pipeline depended on GPU-intensive AI models. Running those models on the same always-on infrastructure would couple application availability with expensive GPU resources.

Decision

I separated the application/API layer from the GPU inference layer.

The application handles requests, workflow state, and storage, while GPU workers are responsible for the actual AI inference.

Why

This gives the two parts of the system different scaling and lifecycle characteristics.

The API does not need a GPU simply because the AI workload does. GPU infrastructure can be scaled according to video-generation demand rather than application traffic.

Trade-off

This introduces additional infrastructure boundaries and makes communication between the application and inference layer more explicit.

Job state, failures, and outputs need to be handled across services.

Evidence

The deployed architecture physically separates AWS application/storage infrastructure from RunPod Serverless GPU inference.

The AI pipeline runs independently from the application server instead of requiring the API itself to host the GPU models.

02

Use on-demand GPU execution

Problem

Video generation is not a continuous workload.

There can be periods where no videos are being generated, followed by periods with multiple GPU-intensive jobs.

Keeping a GPU instance running continuously would mean paying for compute even when the system is idle.

Decision

I used RunPod Serverless for the GPU workloads so that inference capacity could be provisioned around actual jobs.

Why

The infrastructure model better matches the workload:

WorkloadGPU behavior
No jobsCapacity scales down
Video jobGPU execution
More jobsMore execution capacity

This separates application availability from GPU availability.

Trade-off

On-demand execution introduces startup overhead and makes cold-start behavior part of the user experience.

The system has to tolerate the fact that GPU capacity may not always be immediately available.

Evidence

The production deployment uses RunPod Serverless for the GPU-intensive stages rather than keeping a dedicated GPU machine continuously active — GPU compute follows workload demand instead of running permanently.

03

Treat video generation as an asynchronous job

Problem

Generating a lecture video can take significantly longer than a normal HTTP request.

Keeping an API connection open while TTS, video generation, and lipsync execute would make the API responsible for a long-running computation — creating timeout, reliability, and resource-management problems.

Decision

I treated video generation as a job rather than a synchronous API operation.

The API accepts the request and the generation pipeline processes the work independently.

Why

This gives the application a clear separation between the request lifecycle and the generation lifecycle.

The user doesn't need to keep an HTTP request alive for the entire duration of GPU inference.

Trade-off

Asynchronous processing adds state management. The application needs to know whether a job is:

  1. queued
  2. running
  3. completed
  4. failed

It also requires handling retries and partial failures explicitly.

Evidence

The deployed system represents video generation as a processing workflow rather than executing the pipeline inside the API request. The API and GPU execution have independent lifecycles.

04

Break the generation pipeline into explicit stages

Problem

The final lecture video was not produced by one model.

Speech generation, teacher/video generation, visual generation, and synchronization were handled by different components. Treating them as one opaque inference process would make failures hard to identify and components hard to replace.

Decision

I kept the generation process as a sequence of explicit stages.

  1. Input
  2. TTS
  3. Teacher Video
  4. Visual / Slides
  5. Lipsync
  6. Final Video
Why

Each stage has a different computational characteristic and can evolve independently.

A model can be replaced without requiring the entire application architecture to change.

Trade-off

More stages mean more intermediate data, more orchestration, and more potential failure points.

The system becomes more complex than a single inference endpoint.

Evidence

The deployed application integrates VietTTS, MimicTalk, and Stable Diffusion-based processing as separate parts of the workflow — an orchestration of several AI workloads rather than a single model endpoint.

05

Keep storage outside the compute layer

Problem

Generated videos and intermediate artifacts are much larger and longer-lived than the GPU processes that create them.

If outputs were tightly coupled to the GPU execution environment, they could become difficult to access once the compute process finished.

Decision

I treated generated media as persistent application data rather than something owned by the GPU worker.

The compute layer produces the artifact; storage owns its lifecycle.

Why

This allows GPU workers to remain temporary.

A worker can execute a job, produce the output, and disappear without becoming the permanent home of the generated video.

Trade-off

The system needs explicit upload/download paths and additional data movement between compute and storage.

Large video files also make transfer time and storage management part of the system design.

Evidence

The architecture separates GPU execution from persistent video storage and application access — generated artifacts survive independently of the GPU environment.

Outcome

What this deployment taught me

The main deployment challenge wasn't getting the models to run.

It was deciding what should stay alive, what should be temporary, and what should happen asynchronously.

That led to five boundaries:

  1. Application
  2. Job
  3. GPU Execution
  4. Artifact
  5. Storage

Each boundary exists because the workload has a different lifecycle.

The result was a deployment architecture where the application could remain available without keeping expensive GPU infrastructure running continuously, while long-running AI workloads could execute independently.

Engineering focusDeployment architecture, workload isolation, and GPU lifecycle management

Read the build logs for this project

Next · 03 · Optimize AI workflows

Real-time Virtual Human Streaming System