Our AI Models
From concept art to full cinematic video, Layer brings together the industry's leading AI models in one place.
Capabilities
Capabilities
Image
Video
3D
Audio
Sort
Sort by
Google's fast multimodal video model — 720p clips with synchronized native audio from text or a still image.
Reference-guided video generation with up to 5 reference images for consistent characters and style.
Alibaba's Wan 3.0 video model with up to 30s clips, 480p/720p/1080p output, and native audio.
Reference-guided Wan 3.0 video with up to 10 images, 5 videos, and 5 audio files.
MiniMax H3 (Hailuo-03) next-gen open-weight video with native stereo audio at 2K and 24 FPS.
Reference-guided MiniMax H3 Max video with multimodal image, video, and audio inputs at 480p, 768p, or 1080p.
MiniMax H3 Max post-trained video with native stereo audio at 480p, 768p, or 1080p.
MiniMax H3 Max Turbo post-trained video with native stereo audio at 480p, 768p, or 1080p.
Reference-guided MiniMax H3 video with multimodal image, video, and audio inputs at 2K.
Reference-guided Seedance 2.5 video with up to 30 images, 10 videos, and 10 audio files.
Professional-grade video model with cinematic quality, up to 4K output, and synchronized audio generation.
Generate videos guided by reference images, videos, and audio with precise style and character control.
ByteDance's Seedance 2.5 video model with up to 30s clips, 480p/720p/1080p output, and native audio.
Fast, cost-effective variant of Seedance 2 with 720p output and synchronized audio generation.
Fast reference-guided video generation with multi-modal inputs and synchronized audio.
xAI's Grok Imagine 1.5 image-to-video model, animating images into 480p or 720p clips.
Latest PixVerse video model with improved quality and audio generation, supporting up to 1080p resolution.
xAI's video generation model capable of creating high-quality 720p video from text and images.
Extend an existing video with xAI's Grok Imagine, continuing motion from the source ending.
Alibaba's text-to-video and image-to-video model generating 720p or 1080p clips with native audio and multilingual lip-sync.
High-speed video model. Produces rapid, professional-quality 1080p video for fast-paced content.
Alibaba's text-to-video and image-to-video model generating 720p or 1080p clips with native audio and multilingual lip-sync.
Reference-guided video generation with up to 9 reference images for consistent characters and style.
Generate high-quality videos with advanced control. Pro tier with negative prompts.
Generate native 4K videos directly from prompts or images. Cinema-grade output in one step.
Vidu's Q3 Pro video model with up to 16s clips, 720p/1080p output, and native audio.
Generate high-quality videos with prompts, images, or reference elements. Pro tier for premium quality.
Generate native 4K videos with Kling's Omni 3 model. Cinema-grade output in one step.
Generate native 4K videos guided by reference images and elements. Cinema-grade output.
Latest PixVerse video model with enhanced quality and audio generation, supporting up to 1080p resolution.
Speed-optimized PixVerse v5.5 variant for rapid video generation at up to 720p resolution.
Generate videos with prompts, images, or reference elements. Standard tier for balanced quality and speed.
Generate videos with prompts and images. Standard tier with negative prompts.
Runway's high-fidelity image-to-video model with improved prompt adherence and motion quality.
Powerful, user-friendly video model producing high-quality, stylized 1080p video with consistent motion.
Speed-optimized Veo 3.1 model. Delivers high-quality video rapidly for dynamic workflows.
Speed-optimized reference-guided Veo 3.1: character-consistent video from up to three reference images.
Generate videos from images with native audio generation and fluid motion.
Flagship video update: refined control, enhanced visual fidelity, and improved subtle motion details.
Reference-guided Veo 3.1: lock characters and objects across shots using up to three reference images.
Next-generation Seedance video model with improved quality, 1080p output, and integrated audio generation.
Cost-effective Veo 3.1 variant. Balances quality and affordability for high-volume video creation.
Latest standard video model. Improved prompt understanding and visual consistency for daily creation.
Pinnacle of Hailuo T2V series. Cinematic quality with superior coherence, detail, and artistic control.
Robust video model producing crisp 768p video. Offers a solid balance of quality and reliable performance.
Premium video model engineered for professional-grade 1080p output, superior fidelity, and smoother motion.
Multilingual TTS with zero-shot voice cloning and prompt-based voice design.
Powerful video model. Optimized for top-tier 1080p cinematic quality and consistency.
High-quality video model. Tuned for maximum visual fidelity and broadcast-quality 1080p output.
Next-gen video model with enhanced control over narrative, tone, and shot composition. Includes audio.
Speed-optimized LTX 2.5 audio-video model with 720p–4K output and clips up to 20 seconds.
Speed-optimized LTX 2.5 mode that generates video timed to a supplied audio clip.
State-of-the-art multimodal video generation model from Alibaba, with native audio support
Quality-optimized LTX 2.5 audio-video model for high-fidelity 720p and 1080p output.
Quality-optimized LTX 2.5 mode for final visuals synchronized to music, dialogue, or soundtrack.
Google's most controllable TTS model with 200+ audio tags for vocal style and delivery.
Premium 1080p video model. Maximum visual fidelity, capturing intricate details and lifelike expressions.
High-speed T2V variant. Optimized for rapid creation, iteration, and social media content workflows.
Generate new videos guided by prompts, images or videos.
Generate new videos from first and last frame images.
Pika's highest-quality image-to-video model, generating 480p, 720p, or 1080p clips from a still frame.
High-fidelity video model designed for professional-quality results, offering superior detail and nuanced audio.
OpenAI's default GPT Image 2.5 variant: fast, high-quality generation with natural lighting, richer textures, and precise edits up to 4K.
Versatile video model integrating video and audio creation in one seamless, speed-optimized workflow.
Premium video version offering higher resolution (up to 1024p) and enhanced controls for pro projects.
OpenAI's premium GPT Image 2.5 variant: extra precision across edits and 1024px to 4K output, trading longer generation times for tighter control.
Refined Kling model delivering professional-grade 1080p video with improved clarity and motion.
Master-grade video model. Enhanced realism and physics simulation for cinematic, high-impact clips.
OpenAI's latest image model with stronger text rendering, UI generation, and photorealism. Native output up to 4K with three quality tiers.
Next-gen video model generating long, high-fidelity 720p video with unparalleled narrative understanding.
Speed-optimized Veo 3 variant for rapid video creation and iteration, ideal for short-form content.
ElevenLabs' most expressive TTS model with inline audio tags for emotion and delivery.
Speed-optimized LTX video model with extended duration support up to 20 seconds and portrait mode.
Latest LTX video model with sharper details, cleaner audio, and portrait support for professional content creation.
Versatile and efficient video model. Optimized for speed and ideal for short clips and rapid prototypes.
xAI's text-to-speech model with speech tags for expressive delivery in 20+ languages.
Powerful video model for high-fidelity, imaginative content with complex character motion.
Reve's next-generation image model with strong prompt adherence and text rendering, plus native editing.
Reve's remix model - blends one to eight reference images into a single new image, guided by a prompt.
High quality open-source video model from Tencent Hunyuan.
Google's fast, high-quality image generation model with multimodal reasoning.
Meta's Muse Image model with faithful instruction-following and precise multi-reference editing.
Legacy video model generating high-definition, long-form cinematic video content.
Advanced video model. Enhanced visual consistency and detail for high-resolution 1080p content.
Microsoft's photorealistic image model with strong typography and native editing.
GPT Image 1.5 is the latest image generation model from OpenAI, with better instruction following and adherence to prompts.
Top-tier reasoning-first image generation from ByteDance, with built-in online search.
Google's state-of-the-art image generation and editing model.
Microsoft's flagship image model for high-fidelity generation and native editing.
Multilingual text-to-speech with natural voice selection and stability controls.
Google's fastest, most cost-efficient Gemini image generation model.
Runway's next-generation video-to-video model for restyling and transforming existing footage from a prompt.
Runway's video-to-video model for restyling and transforming existing footage from a prompt.
Runway's speed-optimized image-to-video model for rapid, high-quality clip generation.
xAI's image generation model capable of creating high-quality images from text prompts.
The highest-fidelity tier of Luma UNI-1, for hero-quality stills and premium edits.
Fast, high-quality image generation from ByteDance, optimized for creative advertising.
Alibaba WAN 2.7 Pro image model for high-quality text-to-image generation and multi-image editing.
Luma's unified image model for high-fidelity generation and prompt-based editing.
A new-generation image creation model from ByteDance, for both generation and editing.
Unified architecture for image generation and editing. Allows fluid movement from concept to refinement.
Qwen Image 2 Pro tier with higher quality text-to-image generation and enhanced detail.
Accessible video model. Generates engaging 720p clips from text with good visual consistency.
Krea's flagship text-to-image model for high-fidelity generations with distinctive aesthetic range.
State-of-the-art video model. Generates detailed, stylistically diverse clips with fluid motion.
Speed-optimized open-source version of Krea 2 — high-fidelity images in seconds, with the full Krea aesthetic range.
Ideogram's latest text-to-image model with crisp visuals, accurate text, and image-to-image support.
Kling Omni 3 image model. High-quality images with text rendering capabilities up to 4K resolution.
Kling V3 image model. High-quality images with negative prompts, supports up to 2K resolution.
Powerful, versatile OpenAI image model for creative and professional apps.
Latest Qwen image editing model. Supports prompt-guided transformations with enhanced quality.
Qwen Image Edit 2511 with the Multiple Angles LoRA. Re-renders the input image from a chosen camera angle.
Professional design model with extended prompt support (10,000 characters) and advanced style customization.
Impressive open-source video model. Excels at stable, coherent sequences with high visual quality.
Recraft's pro V4 tier, with higher fidelity and finer detail than the standard tier.
Google's image generation and editing model capable of multimodal reasoning.
Professional design model with extended prompt support (10,000 characters) and advanced style customization.
HiDream's full O1 Image model — create, edit, and personalize images up to 2K in one native model.
Professional design model with extended prompt support (10,000 characters) and advanced style customization.
Latest Qwen text-to-image model with enhanced quality, detail, and prompt adherence.
High-speed variant of Ray 2, optimized for rapid video creation. Perfect blend of speed and quality.
Advanced image editing model. Enhanced performance for fine-grained manipulation.
Qwen Image 2 standard text-to-image model with strong prompt adherence and diverse style support.
Large-scale, state-of-the-art video model for stunning realistic and coherent 1080p motion.
Premium model for max editing performance, superior typography, and visual narrative consistency.
Ultra-fast 6B parameter image model from Tongyi-MAI, optimized for near real-time generation.
Bria's one-shot photo restoration, trained on licensed data.
Pro-grade multimodal model for fast, iterative editing, style transfer, and consistency.
Delivers ultra-high-res (up to 4MP) images with superior photorealism, detail, and speed.
Next-gen FLUX model. 6x faster with enhanced prompt adherence and top-tier quality for production.
Flagship commercial T2I model offering superior prompt adherence, quality, and stylistic outputs.
Open-source T2I model. Excels at high-res images from complex text, notable for clear, stylized text.
12B flow transformer fine-tuned with SRPO for exceptional photorealism and polished composition.
HiDream's distilled O1 Image Dev model — create, edit, and personalize images up to 2K in one native model.
Versatile image model for graphic design. Generates legible, stylized text and scalable vector art (SVG).
Open-weight, multimodal model for context-aware image editing. Excels at iterative edits via text.
Latest open-weight MM-DiT model. Major improvements in quality, prompt following, and text rendering.
Open-weight model co-developed with Krea AI. Excels in photorealism and aesthetics.
Powerful, open-weight 12B image model. Excels in image quality, prompt adherence, and commercial use.
Ultra-fast, open-source model. Generates high-quality images in 1-4 steps for rapid prototyping.
ByteDance's SeedVR2 model for high-quality image upscaling.
Professional-grade video upscaling solution utilizing Topaz AI technology for high-quality enhancement.
Enhance image quality with advanced upscaling.
Meshy V6. High-fidelity 3D asset creation focusing on PBR, quad mesh, and face rigging.
Meshy V7. Text, image, and multi-image to 3D with PBR, quad mesh, and optional rigging.
Fast, clean background removal from Ideogram.
Convert raster images to clean, scalable SVG vector graphics with fine-grained control over detail.
Latest 3D model for production-quality assets with superior textures and clean geometry.
Topaz Generative Enhance for enhanced details.
Powerful video upscaling model from ByteDance, optimized for high-quality resolution and fidelity improvement.
Dedicated Meshy 3D tool for remeshing models to optimize topology and reduce polygon count.
Remove video backgrounds for advanced video editing.
Latest Hunyuan 3D model with enhanced quality, multi-view input, and PBR material support.
xAI's Grok Imagine Image 2.0 model for high-quality text-to-image generation and editing.
Topaz's creative image upscaler, reinventing detail with an adjustable creativity dial.
Split a finished image into independently editable layers with Seedream 5.0 Pro Layerize.
Google's multimodal video model — 720p, 1080p, or 4K clips with synchronized native audio from text or a still image.
High-quality image background removal from Pixelcut.
Topaz's manually tuned precision upscaler, biased toward detail.
Very fast upscaling with good quality.
Generate custom sound effects from text descriptions with duration and prompt control.
Topaz Gigapixel precision upscaling, faithful to the original image.
Boost resolution while refining small details and faces.
Sound effects generation and editing with text-to-audio and audio inpainting.
Upscale images with high fidelity or creativity.
Topaz's generative video upscaler, rebuilding detail that the source never had.
High-quality, commercial-use-safe sound effects from text with exact duration control.
Auto-rigs humanoid 3D models and optionally applies a preset animation, returning rigged GLB/FBX.
Topaz's generative video upscaler tuned for archival restoration.
Retextures an existing mesh from a text prompt or reference image, keeping geometry intact.
Segments a 3D mesh into parts, returning a part-grouped model for further editing.
Specialized model to automatically create and sync sound effects (foley) for video content.
Game-ready Tripo model: a single image to clean low-poly 3D with PBR textures.
Rigs a 3D mesh and applies a preset animation from the animation library, returning an animated GLB.
FLUX.3-powered source-faithful video upscaler, up to 4K.
Topaz upscaling that preserves the alpha channel end to end.
Splits a 3D model into parts for editing, printing, or modular asset workflows.
Upscale for a sharper, cleaner, higher-resolution result.
Open-source text-to-audio model for sound effects, field recordings, and instrument samples.
Google DeepMind's next-generation music model with enhanced composition and vocal quality.
Open-source, high-quality 3D model from Microsoft, leveraging a novel field-free sparse voxel structure.
Bria's advanced AI technology designed to increase the resolution of video content efficiently.
Specialized Meshy 3D tool for quickly retexturing imported meshes using text prompts.
Fast, efficient open-source I2V model. Animates still images using an autoregressive approach.
Lightweight, efficient 3D version optimized for less powerful hardware. Delivers good quality assets.
Edit an existing video with FLUX.3 from a text instruction, preserving motion, timing, and framing.
Recraft's pro V4 text-to-vector tier, with higher fidelity than the standard vector tier.
Recraft's V4 text-to-vector model. Generates scalable vector art (SVG) directly from a prompt.
Reference-guided video generation with up to 10 images and 3 short clips for consistent characters and style.
Google Virtual Try-On: dress a person photo with a clothing product image.
FLUX.3-powered video upscaler that adds detail, up to 4K.
Rebuilds an existing mesh's UV layout so it is ready to texture. Geometry is kept as-is and the result is untextured.
Splits a 3D FBX model into separate parts with Hunyuan 3D v3.1.
Topaz SDR-to-HDR conversion for grading and mastering.
Topaz motion deblur, faithful to the source footage.
Topaz's precision video denoiser for pipeline-ready output.
Topaz's video denoiser tuned for extreme noise.
Topaz's half-price video denoiser for high-volume work.
Topaz's high-quality video denoiser, preserving texture.
Topaz film clean-up for dust and scratches.
Topaz's generative restoration, rebuilding natural detail.
Topaz's generative denoiser, rebuilding detail as it cleans.
Topaz's most aggressive conventional denoiser.
Topaz denoising tuned for heavier noise.
Topaz general-purpose denoising at the source resolution.
Topaz's frame interpolator for extreme and complex motion.
Topaz's frame interpolator tuned for real-world motion.
Topaz's general-purpose frame interpolator, retiming footage up to 120fps.
Topaz's softer precision upscaler, tuned to look untouched.
Topaz's maximum-detail precision upscaler.
Topaz's precision upscaler specialised in recovering facial detail.
Topaz's precision video upscaler, enhancing detail without inventing it.
Topaz's restorative upscaler, rebuilding natural detail in damaged images.
Topaz's precision upscaler that denoises and sharpens as it enlarges.
Topaz's animation-tuned precision upscaler, and its cheapest tier.
Topaz's deinterlacing precision upscaler for legacy tape sources.
Topaz's upscaler for extremely low-resolution sources.
Topaz's creative video upscaler, reinventing detail for maximum visual impact.
Topaz's precision upscaler for rendered and CG video.
Topaz's precision upscaler for refining already-clean footage.
Topaz's fastest and cheapest generative video upscaler.
Topaz's creative upscaler with its invented detail biased toward photorealism.
Topaz's precision-leaning generative upscaler.
Topaz's generative image upscaler, rebuilding detail with fewer repeated patterns.
Topaz's detail-preserving precision upscaler for professional photography.
Topaz's highest-quality generative video upscaler.
Topaz's faster generative video upscaler with sharper output.
Topaz's precision upscaler tuned for small, compressed sources.
Topaz's precision upscaler that keeps text and hard shapes crisp.
Topaz's prompt-guided generative upscaler.
Topaz's precision upscaler for rendered art and CG rather than photographs.
Hi3D v3.0 image-to-3D at 2048³ with PBR materials and face counts up to 5M.
Production-quality lipsync that re-articulates a source video to match any new audio track.
Fast, affordable video-to-video lipsync that syncs mouth motion to any audio track.
High-quality realistic lipsync that preserves unique facial details from any new audio track.
sync.so's most powerful lipsync model, re-articulating a source video to any new audio track.
Turns a single still image into a talking character lip-synced to a voice track.
ByteDance's latest OmniHuman, bringing a still photo to life from audio with improved motion and expression.
Retargets a Uthana motion — trained, or generated from a prompt or video — onto your character.
Uthana's automatic character rigging model, preparing an uploaded mesh with a skeleton.
High-performance MiniMax music model for complete songs up to five minutes with structure tags and seed control.
Meshy T2. Text or single image to a compact, cleanly built multi-part mesh with textures, PBR, and optional rigging.
Virtual try-on from Black Forest Labs: dress a person photo with a garment reference.
Meta's SAM 3.1 for fast multi-object image segmentation.
Meta's SAM 3.1 for multi-object video segmentation and tracking.
Frame-synced, licensed music scored from video pacing, mood, and timing.
Synchronized, royalty-free sound effects timed to actions visible in a video.
Mux frame-synced, licensed music onto any video; optionally keep original speech.
Licensed, commercial-use-safe music from a text prompt with exact duration control.
Add synchronized, royalty-free sound effects mixed into the finished video.
Segments a 3D mesh into parts with Hyper3D Bang!
Hyper3D Rodin Gen-2 delivers sharper geometry and cleaner textures from a single image or prompt.
Hyper3D Rodin Gen-2.5 pushes structural detail and surface quality further, with selectable quality.
Hi3D image-to-3D with multi-view support, PBR materials, and configurable face counts.
Textures an existing Hi3D-compatible geometry mesh from a reference image.
Extend an existing video clip with FLUX.3, continuing motion and native audio from the source ending.
Frontier video model from Black Forest Labs with native audio, lipsync, and first/last frame control.
Guide FLUX.3 video generation with up to 10 keyframe images pinned across the clip.
Bytedance Seed Audio TTS with preset voices and zero-shot voice cloning.
Reve's remix model - blends one to eight reference images into a single new image, guided by a prompt.
Recraft's V4.1 text-to-vector model. Generates scalable vector art (SVG) directly from a prompt.
Reve's flagship image model with best-in-class prompt adherence and text rendering, plus native editing.
Recraft's V4.1 Pro text-to-vector model. Generates scalable vector art (SVG) directly from a prompt.
Fast Hunyuan 3D model for rapid 3D generation with PBR material support.
Remove video backgrounds with Pixelcut for clean cutouts.
Bria's v3 video background removal for advanced video editing.
Latest MiniMax music model with native audio-visual generation capabilities.
Updated MiniMax music model with improved vocal quality and arrangement control.
Enterprise-grade audio generation with multi-part compositions and audio inpainting.
Real-time sound effects generation in under 1 second, up to 30 seconds long.
Foundation model for audio source separation using text, visual, or temporal prompts.
Inworld's premium TTS model optimized for expressive game character voices.
Generate multi-speaker dialogue audio from structured text with per-speaker voice control.
Premium tier of Google's Lyria 3 with highest fidelity and compositional complexity.
Transform the voice in an audio recording to a different target voice.
Ultra-fast music generation producing 3-minute tracks in under 10 seconds at 44.1 kHz.
Google DeepMind's high-fidelity 48 kHz music model with fine-grained creative control.
AI music generator producing complete songs with vocals, lyrics, and full instrumentation.
Real-time speech enhancement and noise suppression at 48 kHz full-band audio.
Generate synchronized sound effects, dialogue, and ambient audio from video content.
Isolate clean voice from noisy recordings by removing background noise and music.
Premium AI music generation from ElevenLabs with composition planning and section control.
Fast open-source music generation with lyrics alignment, remixing, and audio editing.
Google's text-to-speech model with multi-speaker support and natural expressiveness.
High-detail Tripo model for production assets with dense geometry and refined textures.
Split an image into layers using Qwen-Image Layered.
Professional-grade 3D model optimized for high-quality, detailed assets with advanced features.
Powerful 3D model focusing on high-quality objects from a single image. Adept at geometry & texture.
Meta SAM 3D Align places body and object meshes into a spatially coherent scene from one image.
Meta SAM 3D Body reconstructs accurate human body shape and pose from a single image.
Meta SAM 3D Objects reconstructs textured 3D geometry from a single real-world image.
Rebuilds a high-poly 3D mesh as a low-poly model at a target face count, baking the original textures onto the result.
Auto-rigs a 3D mesh, returning a rigged GLB. Humanoid meshes come back with a Mixamo-named skeleton for direct use in game engines; other body plans use Tripo's own naming, which Mixamo cannot describe.
Advanced video model bringing a still image of a person to life using audio, producing expressive videos.
Specialized model for creating realistic, audio-driven talking avatars with accurate lip-sync and expressions.
High-quality image background removal.
ESRGAN model for video upscaling, enhancing resolution and detail.
Convert raster images to vector graphics using Recraft.
Speed-optimized 3D generation model designed for rapid prototyping and fast generation times.
Advanced 3D generation model creating high-quality, textured T-pose avatars from a single image.
Highly efficient, open-source I2V model that generates video by predicting the next frame.
Incremental 3D update. Refined performance, improved mesh topology, and texture fidelity.
Powerful, open-source 3D model producing high-res, textured 3D objects from text or image inputs.
Specialized video model. Optimized for a dynamic, live-action feel with naturalistic camera work.
Open-source 3D model creating high-quality objects with realistic materials and geometry from text.