mcp
What this MCP does
Generates and edits AI images, videos, music, and sound effects, and upscales image and video assets.
Tools
Input schema
{'type': 'object', 'properties': {'type': {'enum': ['image', 'video', 'music', 'sound_effect'], 'type': 'string', 'description': 'Which pipeline the job_id belongs to. REQUIRED as "music" for music jobs and "sound_effect" for sound-effect jobs (their numeric ids would otherwise be read as image jobs). Optional otherwise: uuids are detected as video, plain numbers default to image.'}, 'job_id': {'type': 'string', 'description': 'Optional job ID to check â\x80\x94 a number for image/music/sound-effect jobs, a uuid for video jobs. If omitted, returns all active generation jobs (images, videos, music, and sound effects).'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['media'], 'properties': {'genre': {'enum': ['Pop', 'Rock', 'Hip-Hop & Rap', 'R&B & Soul', 'Electronic & Dance', 'Lo-fi & Chill', 'Ambient', 'Jazz', 'Classical', 'Country', 'Folk & Acoustic', 'Metal', 'Blues', 'Latin', 'K-Pop & J-Pop', 'Reggae', 'Soundtrack & Cinematic', 'Other'], 'type': 'string', 'description': "Audio posts only: the track's genre â\x80\x94 EXACTLY one of these values. Omit to have it classified automatically after publishing."}, 'media': {'type': 'array', 'items': {'type': 'object', 'required': ['file'], 'properties': {'file': {'type': 'string', 'description': 'The media file â\x80\x94 a URL (generation result link, upload_media URL, or any public URL), a data URI, or raw base64.'}, 'model': {'type': 'string', 'description': 'Optional generation info: the model that made this image (images only).'}, 'prompt': {'type': 'string', 'description': 'Optional generation info: the prompt that made this image (images only).'}}, 'additionalProperties': False}, 'maxItems': 6, 'minItems': 1, 'description': "1â\x80\x936 media items. Rules: 1â\x80\x936 images, OR 1 video plus up to 3 images, OR exactly 1 audio track (audio can't mix)."}, 'lyrics': {'type': 'string', 'description': 'Audio posts only: lyrics to show with the track.'}, 'content': {'type': 'string', 'description': "The post's text/caption (optional â\x80\x94 media-only posts are fine)."}, 'theme_id': {'type': 'integer', 'minimum': 1, 'description': "Optional: enter the post in a live clan theme â\x80\x94 the id from list_my_themes (or the number in a budgetpixel.com/themes/<id> link). Works with any media type; posting to a theme of a clan the user hasn't joined joins them to it, like on the site."}, 'song_name': {'type': 'string', 'description': "Audio posts only: the track's display name."}, 'cover_image': {'type': 'string', 'description': 'Audio posts only: cover art â\x80\x94 an image URL (generation link / upload_media URL), data URI, or base64 (max 15MB). Goes in THIS field, never as a media item.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['prompt'], 'properties': {'seed': {'type': 'integer', 'description': 'Seed for reproducible generation'}, 'model': {'type': 'string', 'description': 'AI model id. Omit for the default (seedream-5.0-pro â\x80\x94 premium flagship: top quality, strongest identity preservation, great for edits). For fast & cheap drafts pick flux-2-klein (free on Pro/Ultra plans) or seedream-5.0-flash (flat 25 credits, free references). Other strong picks: gpt-image-2.5-sunburst (OpenAI precision edits, premium quality), gpt-image-2.5-flare (fast OpenAI everyday model), seedream-4.5 (photorealism), gpt-image-2 (text in images), nano-banana-2 (premium editing), qwen-image-3.0 (all-rounder). Call list_models for the full catalog with prices and input limits.'}, 'prompt': {'type': 'string', 'description': 'Text prompt describing the image to generate'}, 'quality': {'type': 'string', 'description': 'Optional rendering quality for models that support it (gpt-image-2: low/medium/high; gpt-image-2.5-sunburst and gpt-image-2.5-flare: low/medium/high/xhigh/max; grok-imagine-image-2: low/medium â\x80\x94 no high tier). Higher costs more. On grok-imagine-image-2 quality x resolution sets the price (low 1K=50, low 2K=80, medium 1K=80, medium 2K=100 credits). Omit for default.'}, 'background': {'enum': ['opaque', 'transparent'], 'type': 'string', 'description': 'Optional output background for models that support it (seedream-5.0-flash, seedream-5.0-pro â\x80\x94 see list_models). "transparent" keeps a transparent background when EDITING: pass exactly one input image, a PNG that already has transparent pixels (a cutout, sticker or logo), and you get back a PNG with transparency. It cannot make a transparent image from text alone. Omit (or "opaque") for a normal image.'}, 'num_images': {'type': 'integer', 'maximum': 4, 'minimum': 1, 'description': 'Number of images to generate (1-4)'}, 'resolution': {'type': 'string', 'description': 'Optional output size (per-model options, e.g. seedream-5.0-pro & seedream-5.0-flash: 1K/2K; gpt-image-2 & nano-banana-2: up to 4K â\x80\x94 see list_models). Omit for the model default. Larger costs more.'}, 'aspect_ratio': {'enum': ['1:1', '4:3', '3:4', '16:9', '9:16', '3:2', '2:3', '21:9'], 'type': 'string', 'description': 'Aspect ratio of the generated image'}, 'input_images': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Optional input image(s) for image-to-image / editing / multi-reference â\x80\x94 each a public image URL, a data URI, or raw base64. Per-model support: some models take 1, some several, some none (see max_input_images in list_models). Omit for text-to-image. To edit an image the user just generated, pass its URL here.'}, 'negative_prompt': {'type': 'string', 'description': 'What to avoid in the generated image'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {'model': {'type': 'string', 'description': 'Music model id. Omit for the default (music-3.0 â\x80\x94 latest MiniMax, all-purpose songs, 200 credits; passing a video auto-selects sonilo-video-music). Alternatives: music-2.6 (previous generation, same price), mureka-v9 (sings exact lyrics, cheapest at 60), lyria-3 (high-fidelity, image conditioning, 100), sonilo-music (instrumental/background music of an exact length, priced per second â\x80\x94 pass duration). Call list_models for details.'}, 'video': {'type': 'string', 'description': "Video-to-music: URL of a video (â\x89¤360 s, â\x89¤30 MB for public URLs) to compose a licensed soundtrack FOR â\x80\x94 the music follows the clip's content, pacing and mood, and matches its length. Must be a URL â\x80\x94 upload local files with upload_media first. Uses sonilo-video-music: 15 credits per second of the video's measured duration (10-second minimum). prompt becomes optional direction (genre/mood/instrumentation); no lyrics or duration."}, 'format': {'type': 'string', 'description': 'music-2.6/music-3.0 only: output audio format, wav (default) or mp3.'}, 'images': {'type': 'array', 'items': {'type': 'string'}, 'description': 'lyria-3 only: up to 10 reference images to condition the mood â\x80\x94 each a URL, data URI, or raw base64 (upload local files with upload_media).'}, 'lyrics': {'type': 'string', 'description': 'Lyrics to sing. mureka-v9 sings them VERBATIM ([verse]/[chorus] tags supported); music-2.6/music-3.0 polish them unless lyrics_optimizer is false; lyria-3 treats them as optional guidance.'}, 'prompt': {'type': 'string', 'description': 'Style/mood description of the track. Required for lyria-3 (10â\x80\x932000 chars) and for instrumental tracks; optional on music-2.6/music-3.0 when lyrics are given.'}, 'duration': {'type': 'integer', 'maximum': 360, 'minimum': 5, 'description': 'sonilo-music only: track length in seconds (5â\x80\x93360, default 30). Billed per second with a 10-second minimum (4 credits/s â\x86\x92 30 s = 120 credits). Other music models ignore length (sonilo-video-music sizes the music to the video).'}, 'instrumental': {'type': 'boolean', 'description': 'Generate an instrumental track (no vocals). music-2.6/music-3.0 and mureka-v9 (mureka requires a prompt in this mode).'}, 'vocal_gender': {'enum': ['auto', 'female', 'male'], 'type': 'string', 'description': 'mureka-v9 only: preferred vocal gender for songs (default auto).'}, 'lyrics_optimizer': {'type': 'boolean', 'description': 'music-2.6/music-3.0 only. Write/polish lyrics from the prompt â\x80\x94 DEFAULT TRUE (matches web/API). Set false to sing the provided lyrics as-is.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {'model': {'type': 'string', 'description': 'Sound-effect model id. Omit for the default (sonilo-sfx for text; sonilo-video-sfx is selected automatically when video is given). Call list_models for details.'}, 'video': {'type': 'string', 'description': "Optional INPUT VIDEO URL to sync the effects to (frame-accurate Foley/ambience/impacts for the clip's visible events). Must be a URL â\x80\x94 upload local files with upload_media first; public URLs OK (â\x89¤30 MB, â\x89¤360 s). Uses sonilo-video-sfx: 15 credits per second of the video (3-second minimum); returns the SFX track AND the video with the effects mixed in. prompt becomes optional direction."}, 'format': {'enum': ['mp3', 'wav'], 'type': 'string', 'description': 'Output audio format: mp3 (default, small) or wav (lossless, for editing).'}, 'prompt': {'type': 'string', 'description': "Description of the sound effect â\x80\x94 Foley, ambience, UI sounds, transitions, impacts, action (e.g. 'heavy rain on a tin roof with distant thunder'). Up to 1000 characters."}, 'duration': {'type': 'integer', 'maximum': 180, 'minimum': 1, 'description': 'Text mode: length of the clip in seconds (1â\x80\x93180, default 10). Billed per second with a 3-second minimum â\x80\x94 10 s = 50 credits on the default model. Ignored when a video is given.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['prompt'], 'properties': {'audio': {'type': 'string', 'description': 'Optional reference audio URL (MP3/WAV/OGG; upload local files with upload_media). Must accompany an image, reference_images, or video input. Not for minimax-h3 / seedance-2.5 / wan-3.0-video(-prime) â\x80\x94 use reference_audios there.'}, 'image': {'type': 'string', 'description': "Optional START FRAME for image-to-video â\x80\x94 a public image URL, a data URI, or raw base64. Can't be combined with reference_images or video."}, 'model': {'type': 'string', 'description': 'Video model id. Omit for the default (seedance-2.5 â\x80\x94 flagship multimodal: text-to-video, image-to-video, reference images/videos/audio; pass a clip as reference_videos to edit/extend; 480p/720p/1080p, up to 30s). wan-3.0-video is the Alibaba all-in-one alternative â\x80\x94 same modes, up to 30s, and ALL input media free (480p 60 / 720p 120 / 1080p 240 credits/sec output-only; wan-3.0-video-prime is the same model generating much faster at 85/170/340). For cheaper drafts pick seedance-2.0-mini (480p 60 / 720p 120 credits/sec). Call list_models for per-resolution prices.'}, 'video': {'type': 'string', 'description': "Optional INPUT VIDEO for video-to-video editing (2-15s) â\x80\x94 must be a URL (upload local files with upload_media first). Can't be combined with image/end_image. On the seedance-2.0 family the INPUT seconds are billed at half the output per-second rate on top of the output. On minimax-h3-max the video is EXTENDED instead (1.6-60s source; length_seconds = new footage appended; source billed 170 credits/sec, first 15s)."}, 'prompt': {'type': 'string', 'description': 'Text prompt describing the video to generate (or the edit to apply, when passing an input video)'}, 'end_image': {'type': 'string', 'description': 'Optional END FRAME used together with image (the start frame) â\x80\x94 the video interpolates between the two.'}, 'resolution': {'type': 'string', 'description': 'Output resolution â\x80\x94 PRICING VARIES A LOT (seedance-2.5: 480p=150, 720p=330, 1080p=750 credits per SECOND, default 720p; seedance-2.0: 1080p=550/4k=1250). Prefer 480p unless asked.'}, 'aspect_ratio': {'type': 'string', 'description': "Aspect ratio (seedance-2.5: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 â\x80\x94 default 16:9). On minimax-h3 it applies to TEXT-TO-VIDEO only. On seedance-2.5 it applies to t2v AND jobs with reference images/audio; a start/end frame or a reference VIDEO forces adaptive (output follows the input's frame). On wan-3.0-video (and -prime) it is honored in EVERY mode (even with a first frame or reference media) and also accepts 'adaptive' (its default â\x80\x94 the model picks a suitable ratio from the inputs)."}, 'generate_audio': {'type': 'boolean', 'description': 'Whether the model should generate audio with the video (no price impact on seedance-2.5/2.0).'}, 'negative_prompt': {'type': 'string', 'description': 'What to avoid in the generated video'}, 'duration_seconds': {'type': 'integer', 'description': 'Video length in seconds (seedance-2.5: 4â\x80\x9330, default 5; seedance-2.0: 4â\x80\x9315). Cost scales linearly.'}, 'reference_audios': {'type': 'array', 'items': {'type': 'string'}, 'description': "minimax-h3 / minimax-h3-max / seedance-2.5 / wan-3.0-video(-prime) only: reference AUDIO clip URLs (WAV/MP3; h3 â\x89¤3/15s combined, 2.5 â\x89¤5/30s combined, wan-3.0(-prime) â\x89¤5/15s combined; free; upload local files with upload_media). Can't be combined with image/end_image."}, 'reference_images': {'type': 'array', 'items': {'type': 'string'}, 'description': "Optional reference images guiding the video (subject consistency) â\x80\x94 each a URL, data URI, or base64. seedance-2.5: up to 15; wan-3.0-video / wan-3.0-video-prime: up to 10 (free); seedance-2.0: up to 9 alone, 6 alongside an input video. Can't be combined with image/end_image."}, 'reference_videos': {'type': 'array', 'items': {'type': 'string'}, 'description': "minimax-h3 / minimax-h3-max / seedance-2.5 / wan-3.0-video(-prime) only: reference VIDEO clip URLs guiding motion/identity â\x80\x94 on seedance-2.5 and wan-3.0-video(-prime) also how you edit/extend a clip (h3: â\x89¤3, 15s combined; 2.5: â\x89¤5, 30s combined; wan-3.0(-prime): â\x89¤5, 15s combined, and input + output â\x89¤ 30s; upload local files with upload_media). h3 and 2.5 bill INPUT seconds on top of the output (h3 160/s; 2.5 80/160/375 per s at 480p/720p/1080p); wan-3.0-video(-prime) input is FREE. Can't be combined with image/end_image."}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {'type': {'enum': ['image', 'video', 'music', 'sound_effect'], 'type': 'string', 'description': 'Which history to fetch: image (default), video, music, or sound_effect.'}, 'limit': {'type': 'integer', 'maximum': 50, 'minimum': 1, 'description': 'Number of recent generations to return (1-50, default 10)'}, 'offset': {'type': 'integer', 'minimum': 0, 'description': 'Offset for pagination (default 0)'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['job_id'], 'properties': {'type': {'enum': ['image', 'video', 'music', 'sound_effect'], 'type': 'string', 'description': 'What kind of generation job_id is: image (default), video, music, or sound_effect.'}, 'job_id': {'type': 'string', 'description': 'The generation to keep â\x80\x94 the job_id from generate_image / generate_music / generate_sound_effect / check_generation_status / get_generation_history (numeric), or the video job uuid from generate_video.'}, 'folder_id': {'type': 'integer', 'description': 'Optional asset folder id to save into (defaults to Uncategorized).'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'properties': {'sort': {'enum': ['popular', 'newest', 'top'], 'type': 'string', 'description': 'popular (default â\x80\x94 best matches first when there is a query), newest, or top (highest rated).'}, 'type': {'enum': ['photo', 'illustration'], 'type': 'string', 'description': 'Optional: photos only, or illustrations (including 3D and painted art) only.'}, 'color': {'enum': ['red', 'orange', 'yellow', 'green', 'turquoise', 'blue', 'purple', 'pink', 'white', 'gray', 'black', 'brown'], 'type': 'string', 'description': 'Optional dominant-colour filter.'}, 'limit': {'type': 'integer', 'maximum': 20, 'minimum': 1, 'description': 'How many results to return (default 10, max 20). The price is the same for any limit.'}, 'query': {'type': 'string', 'description': 'What the image should show, in plain words â\x80\x94 e.g. "misty pine forest at sunrise", "latte art on a wooden table", "abstract purple gradient background". Max 200 characters. Optional only when browsing a category.'}, 'offset': {'type': 'integer', 'maximum': 5000, 'minimum': 0, 'description': 'For the next page: the previous offset plus the number returned. Each page is a new 10-credit search.'}, 'category': {'enum': ['nature', 'backgrounds', 'wallpapers', 'textures', 'animals', 'food', 'travel', 'transportation', 'architecture', 'business', 'industry', 'technology', 'people', 'education', 'music', 'art', 'abstract', 'holidays', 'science', 'health', 'sports'], 'type': 'string', 'description': 'Optional category filter.'}, 'orientation': {'enum': ['landscape', 'portrait', 'square'], 'type': 'string', 'description': 'Optional shape filter: landscape for headers and slides, portrait for phone wallpapers and stories.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['file'], 'properties': {'file': {'type': 'string', 'description': 'The file to upload, as raw base64 or a data URI (data:image/png;base64,...). Read the local file and pass its base64 here. Max 50MB. Images: PNG/JPEG/WebP/GIF · Video: MP4/WebM · Audio: MP3/WAV/OGG.'}}, 'additionalProperties': False}
Input schema
{'type': 'object', 'required': ['image'], 'properties': {'image': {'type': 'string', 'description': "The image to upscale: an http(s) URL (a generated image's URL, or one from upload_media for local files), a data URI, or raw base64. Max ~20MB."}, 'model': {'type': 'string', 'description': 'Upscale model: p-image-upscale (default â\x80\x94 Pruna AI, results in seconds, priced by target resolution 4/8/16/32 MP = 10/15/30/60 credits) or clarity-upscaler (creative detail enhancement, flat 30 credits, ~60s).'}, 'prompt': {'type': 'string', 'description': 'clarity-upscaler only: optional guiding prompt describing the image content.'}, 'target': {'enum': [4, 8, 16, 32], 'type': 'integer', 'description': 'p-image-upscale only: target output resolution in megapixels (sets the price: 4=10, 8=15, 16=30, 32=60 credits; default 4).'}, 'creativity': {'type': 'number', 'description': 'clarity-upscaler only: how much new detail the model may invent (0.3-0.9, default 0.35).'}, 'enhance_details': {'type': 'boolean', 'description': 'p-image-upscale only: enhance textures and fine detail (may increase contrast).'}, 'enhance_realism': {'type': 'boolean', 'description': 'p-image-upscale only: improve realism â\x80\x94 recommended for AI-generated images.'}}}
Input schema
{'type': 'object', 'properties': {'video': {'type': 'string', 'description': "The video to upscale: an http(s) URL (a generated video's URL, or one from upload_media for local files). Up to 20 seconds, â\x89¤100MB. Required unless job_id is set."}, 'job_id': {'type': 'integer', 'description': 'A previously returned upscale job id â\x80\x94 pass it alone to check progress and fetch the finished video URL.'}, 'target_fps': {'enum': [30, 60], 'type': 'integer', 'description': 'Target frame rate (default 30). 60 fps doubles the per-second rate.'}, 'target_resolution': {'enum': ['1080p', '4k'], 'type': 'string', 'description': 'Target output resolution (default 1080p). 4K costs 4x: 1080p = 15/30 credits per input second (30/60fps), 4K = 60/120.'}}}
Recent tool changes
Similar MCP servers
MiOffice — AI-Powered Workspace Studio
Provides browser-based tools for processing PDFs, images, video, and audio, including generation, enhancement, conversion, transc…
BlitzReels Video Editor
Creates and edits short-form videos, including timelines, captions, transitions, AI-generated visuals, music, voiceovers, sound e…
Morpha
Enables creating, animating, and exporting layered short-form video projects.
switch
Enables management and exploration of an account-scoped library of AI-generated images and videos.
framesail
Creates long-form YouTube videos through script generation, storyboarding, visual assets, voiceover, music, scene composition, an…
Magnific
Provides image enhancement and design tools, AI image and audio generation, text-to-speech, creation management, and reusable cre…
Uwear
Supports AI fashion photoshoot production with garments, outfits, avatars, locations, art direction, image generation, backdrops,…
Creative Claw
Generates and edits branded images, video, audio, speech, and 3D-related creative assets, with media processing and reusable them…