此 MCP 可以做什么
Enables management and exploration of an account-scoped library of AI-generated images and videos.
工具
输入模式
{'type': 'object', 'required': ['project_id', 'asset_ids'], 'properties': {'asset_ids': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Up to 12 owned image ids from your Switch library.'}, 'project_id': {'type': 'string'}}}
输入模式
{'type': 'object', 'required': ['video_url'], 'properties': {'question': {'type': 'string', 'description': 'Optional. What to find out about the video — tone, structure, on-screen text, sentiment, etc.'}, 'video_url': {'type': 'string', 'description': 'A public https video URL (YouTube ok), OR one of your own Switch videos — a video/asset id, or the download_url / view_url from list_my_videos or get_video_status. Switch resolves its own links to the file for you.'}}}
输入模式
{'type': 'object', 'required': ['video_url'], 'properties': {'force': {'type': 'boolean', 'description': 'Optional. Re-running the same video and question returns the existing report without charging again; pass true to force a fresh, freshly billed analysis.'}, 'question': {'type': 'string', 'description': 'Optional. Something to pay special attention to.'}, 'video_url': {'type': 'string', 'description': 'A public https video URL (YouTube ok), OR one of your own Switch video ids.'}, 'duration_seconds': {'type': 'number', 'description': 'Length in seconds. Required for external files; YouTube and your own Switch videos are measured automatically.'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['warm_golden', 'cool_noir', 'moody_desaturated'], 'type': 'string', 'description': 'warm_golden = late-afternoon honey. cool_noir = neon-fill desaturated. moody_desaturated = soft window low-contrast.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['classic', 'golden_hour'], 'type': 'string', 'description': 'classic = Hasselblad H6D studio. golden_hour = Canon R5 outdoor.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['testino_glossy', 'klein_dark_cinematic', 'inez_vinoodh_hard_flash', 'leibovitz_painterly', 'walker_dreamlike', 'lindbergh_bw_natural', 'cass_bird_off_duty'], 'type': 'string', 'description': 'Photographer attribution drives the lighting + camera + grade stack.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['digital_phone', 'film_pointshoot', 'off_duty_intimate'], 'type': 'string', 'description': 'digital_phone = Sony A7IV + 50mm f/1.4 GM phone-style realism. film_pointshoot = Contax T2 35mm Portra 400. off_duty_intimate = Cass Bird natural-window editorial.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['canon_85mm', 'hasselblad_80mm'], 'type': 'string', 'description': 'canon_85mm = Canon R5 portrait standard. hasselblad_80mm = medium-format luxury.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['neon_noir_action', 'glamour_finance_excess', 'superhero_blockbuster', 'video_game_character', 'generic_action_thriller'], 'type': 'string', 'description': 'neon_noir_action = wet streets + neon + anamorphic. glamour_finance_excess = 1980s Wall Street mahogany / gold. superhero_blockbuster = comic-book key art. video_game_character = Unreal-Engine character render. generic_action_thriller = ARRI cinematic.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['clean_studio', 'lifestyle', 'macro_detail', 'flat_lay'], 'type': 'string', 'description': 'clean_studio = seamless backdrop hero. lifestyle = product in use. macro_detail = extreme close-up texture. flat_lay = top-down catalog.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['hotel_hero', 'rural_property', 'scenic_view', 'drone_aerial', 'lifestyle', 'interior'], 'type': 'string', 'description': 'hotel_hero = property is the star. rural_property = country estate. scenic_view = pure landscape. drone_aerial = top-down or 45° from above. lifestyle = model + destination. interior = inside the property.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['phone_shot', 'film_pointshoot', 'mirror_selfie', 'car_selfie'], 'type': 'string', 'description': 'phone_shot = iPhone-style snap. film_pointshoot = Contax T2 grain. mirror_selfie = bathroom/bedroom mirror. car_selfie = inside-the-car phone.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['style'], 'properties': {'style': {'enum': ['warm_amber_tropical', 'hanalei_cinematic', 'high_key_cyan_beach'], 'type': 'string', 'description': 'warm_amber_tropical = warm honey grade with golden haze. hanalei_cinematic = soft golden mist + infinity pool reflection. high_key_cyan_beach = bright daylit cyan ocean.'}, 'subject': {'type': 'string', 'description': 'What you want to shoot. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony".'}}}
输入模式
{'type': 'object', 'required': ['run_id', 'quote_id', 'approved'], 'properties': {'run_id': {'type': 'string'}, 'approved': {'type': 'boolean'}, 'quote_id': {'type': 'string'}}}
输入模式
{'type': 'object', 'required': ['message'], 'properties': {'message': {'type': 'string', 'description': 'What to say to the operator. Plain words. Name staged Studio references as @Image1 etc.'}, 'wait_seconds': {'type': 'number', 'description': 'How long to wait for words before returning (5 to 55, default 45). If it is still working, call again with the same conversation_id.'}, 'conversation_id': {'type': 'string', 'description': 'The conversation_id from an earlier answer to continue that conversation; omit to start a new one.'}}}
输入模式
{'type': 'object', 'required': ['run_id'], 'properties': {'run_id': {'type': 'string'}}}
输入模式
{'type': 'object', 'required': ['taskId'], 'properties': {'taskId': {'type': 'string', 'description': 'Task id from generate_image or list_my_tasks.'}}}
输入模式
{'type': 'object', 'required': ['project_id', 'state'], 'properties': {'state': {'enum': ['default', 'media_assets', 'auto_clean', 'editor_project', 'editor_design_basic', 'editor_design_hsl', 'editor_design_vintage', 'editor_design_curves', 'editor_design_color_wheel', 'editor_design_mask', 'editor_layers', 'editor_audio', 'editor_renders', 'editor_variables', 'editor_history', 'editor_comments', 'editor_text_basic', 'editor_text_effects', 'editor_captions_auto', 'editor_captions_templates', 'editor_captions_packaging', 'editor_captions_lyrics', 'editor_captions_manual', 'editor_captions_emojis', 'web', 'workspace_storyboard', 'workspace_code', 'workspace_comps', 'record', 'captions_menu', 'keyboard_shortcuts', 'export', 'layout', 'preview_social', 'preview_quality', 'timeline_linkage', 'timeline_preview_options', 'timeline_transition', 'timeline_add_track', 'help', 'workflow_screen_record'], 'type': 'string', 'description': 'The real Editor panel or menu state to open before capture.'}, 'fill_tray': {'type': 'boolean', 'description': 'real-tray-assets: before capturing any state, put the newest real picture and clip from your library in the project tray if it has none. Always on for workflow_screen_record.'}, 'project_id': {'type': 'string'}, 'prompt_text': {'type': 'string', 'description': 'workflow_screen_record only. The words typed into the prompt box on camera, up to 160 characters.'}, 'record_seconds': {'type': 'number', 'description': 'workflow_screen_record only. How long the cursor rests on Send at the end, 3 to 12 seconds, default 6.'}, 'include_controls': {'type': 'boolean', 'description': 'Default true. Also capture every rendered interactive control, including controls reached by scrolling, as a named asset.'}}}
输入模式
{'type': 'object', 'properties': {'estimatedCost': {'type': 'number', 'description': 'Optional dollar amount to test against your daily limit.'}}}
输入模式
{'type': 'object', 'required': ['taskId'], 'properties': {'taskId': {'type': 'string', 'description': 'Task id to check.'}}}
输入模式
{'type': 'object', 'required': ['video_url'], 'properties': {'video_url': {'type': 'string', 'description': 'A public https video URL, OR one of your own Switch video ids.'}, 'duration_seconds': {'type': 'number', 'description': 'Clip length in seconds. Required for external URLs; your own Switch videos are measured automatically.'}}}
输入模式
{'type': 'object', 'properties': {'name': {'type': 'string', 'description': 'Project name.'}}}
输入模式
{'type': 'object', 'required': ['action_id', 'decision', 'card_label', 'card_detail'], 'properties': {'decision': {'enum': ['allow', 'deny'], 'type': 'string'}, 'action_id': {'type': 'string', 'description': 'pending_action.action_id from the last answer.'}, 'card_label': {'type': 'string', 'description': 'Repeat pending_action.label exactly as you showed it to the creator; the decision is refused without it.'}, 'card_detail': {'type': 'string', 'description': 'Repeat the first 40 characters of one detail value from pending_action.details (the command, file or prompt) as you showed it; refused if it does not belong to this card.'}, 'wait_seconds': {'type': 'number', 'description': "How long to wait for the operator's next words (5 to 55, default 45)."}}}
输入模式
{'type': 'object', 'required': ['project_id'], 'properties': {'project_id': {'type': 'string'}}}
输入模式
{'type': 'object', 'properties': {'mode': {'enum': ['text-to-video', 'image-to-video', 'reference-to-video', 'frame-to-frame', 'motion', 'omni', 'video-edit', 'upscale'], 'type': 'string', 'description': 'Video mode. Must be supported by the chosen model (see list_video_models).'}, 'audio': {'type': 'boolean', 'description': 'Generate audio. ON by default on Seedance 2.5 (text, image and reference), on Omni, and on every WAN 3.0 lane (text, image, reference); set false for a silent clip. Models without audio ignore this. See list_video_models for which models generate audio and the max seconds with vs without audio.'}, 'model': {'type': 'string', 'description': "Model id from list_video_models (e.g. kling-v3, seedance-2.0-t2v, wan-3.0-t2v, h3-max-t2v, topaz). Or prefer option_id from list_video_models. MiniMax H3 Max ids: h3-max-t2v, h3-max-i2v, h3-max-r2v (option ids h3max-text / h3max-image / h3max-reference). MiniMax H3 Max Turbo (preview, about twice as fast at half the price, no reference lane): h3-max-turbo-t2v, h3-max-turbo-i2v (option ids h3max-turbo-text / h3max-turbo-image). WAN 3.0 ids: wan-3.0-t2v, wan-3.0-i2v, wan-3.0-r2v (option ids wan30-text / wan30-image / wan30-reference): 720p or 1080p, any whole second 2 to 30, generated audio on by default, image lane takes image_url and an optional end_image_url, reference lane takes up to 10 reference_image_urls and 5 reference_video_urls. Switch Cinema (option ids switch-cinema-480p / switch-cinema-720p, model switch-cinema-motion): remakes the clip in reference_video_urls with the person in 1 to 8 reference_image_urls, keeps the clip's motion, no duration setting, priced per second of the clip you attach (8 tokens at 480p, 17 at 720p)."}, 'subject': {'type': 'string', 'description': 'The shot: subject + motion + scene (video needs motion language, e.g. "slow push-in"). Name the creator\'s staged Studio references as @Image1, @Video1 (the Video tab strip, then the image strip; see get_my_active_references) and they attach automatically; nothing attaches unless named. WAN 3.0: refer to references by their place ("Image 1", "Video 1"), not by @tags; there is no negative-prompt field, so write exclusions inline as "Avoid: ...". MiniMax H3 Max: put spoken words in double quotes and the character says them lip-synced; name references by place (Image 1, Video 1, Audio 1).'}, 'duration': {'type': 'string', 'description': 'Clip length in seconds, or "auto" to let the model choose. Seedance 2.5 does 4-30s (auto bills the 30s cap up front and refunds the unused seconds); Seedance 2.0 does 4-15s; Switch Video (WAN 2.7) does 5/10/15; WAN 3.0 takes any whole second from 2 to 30; Kling/Switch Video Edit cap at 10; MiniMax H3 Max takes whole seconds 5 to 15 (no auto) — see each model\'s durations in list_video_models.'}, 'image_url': {'type': 'string', 'description': 'Required for image-to-video / frame-to-frame / motion. Accepts EITHER a Switch asset id (from show_media / list_my_assets / upload_media) OR a public https url. An asset id is resolved server-side, so just pass the id you have — no need to fetch a url first.'}, 'option_id': {'type': 'string', 'description': 'Catalog id from list_video_models. Native Seedance Draft: seedance-25-draft-text, seedance-25-draft-image, seedance-25-draft-reference (480p with audio). Make Full HD: seedance-25-draft-final with the completed draft job_id in video_url, no new creative settings. Quote then generate within 7 days.'}, 'task_type': {'enum': ['auto', 'edit', 'extend'], 'type': 'string', 'description': "Seedance 2.5 reference mode only: declare what you are doing with the reference clip so the size and length rules are checked immediately instead of failing a minute in. auto = a new take from the references (default), edit = change something inside the clip, extend = continue the clip. Editing and extension inherit the source clip's size, and editing also inherits its length."}, 'video_url': {'type': 'string', 'description': 'Required for video-edit and upscale (the source clip). Accepts one of YOUR Switch videos — a job id from list_my_videos / get_video_status, or its download_url / view_url — or any publicly downloadable https URL. Switch resolves its own videos for you; no need to scrape a page for the file.'}, 'resolution': {'enum': ['480p', '720p', '768p', '1080p', '4k'], 'type': 'string', 'description': 'Output resolution. Defaults to 1080p where the model supports it. 720p is cheaper and faster. 480p is the cheapest. Seedance 2.5 offers 480p, 720p and 1080p on every mode (no 2K/4K). 4K is only on Kling v3 text/image and Kling Omni; Seedance text-to-video is 720p only. MiniMax H3 Max offers 480p and 768p only (its reference lane is 768p only). Each model lists its available resolutions in list_video_models.'}, 'aspect_ratio': {'type': 'string', 'description': 'e.g. 9:16, 16:9, 1:1. Must be allowed for the model (see list_video_models).'}, 'end_image_url': {'type': 'string', 'description': 'End frame for frame-to-frame mode.'}, 'quoted_credits': {'type': 'number', 'description': 'Exact quoted_credits returned by quote_video; used with quote_fingerprint.'}, 'quote_fingerprint': {'type': 'string', 'description': 'From quote_video. Pass with quoted_credits to reject a changed quote before charging.'}, 'face_reference_ids': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Face reference asset ids from upload_reference_asset (frame_type "face") — the ONLY way to use a face/likeness reference in video. Each id is verified server-side (your own untouched original + identity verification) before the shot fires or is charged; URLs and generic uploads here are rejected.'}, 'enhancement_receipt': {'type': 'string', 'description': 'Receipt from enhance_video_prompt. Pass its enhanced_prompt as subject with the same model, settings and references. Prevents duplicate enhancement. Changed/expired receipts fail before generation.'}, 'reference_audio_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Seedance reference/omni only: up to 3 reference audio files to drive synthesized audio. Requires at least one reference image or video. MiniMax H3 Max reference (h3max-reference) also takes up to 3 audio files as conditioning (voice/sound guidance); it generates native audio on every clip regardless.'}, 'reference_image_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'GENERIC reference images (products, scenery, outfits, style). Each entry accepts EITHER a Switch asset id (from show_media / list_my_assets / upload_media / get_my_active_references) OR a public https url — asset ids are resolved server-side. Seedance reference/omni accepts up to 9; Kling Omni up to 7. WAN 3.0 reference (option wan30-reference, model wan-3.0-r2v) accepts up to 10; name each one by its place in the prompt ("the woman in Image 1", "the jacket from Image 3"). For Seedance, at least one image or video reference is required. For a person\'s face/likeness use face_reference_ids instead. Switch Cinema (switch-cinema-480p / switch-cinema-720p): 1 to 8 pictures of the person to put into the clip, required.'}, 'reference_video_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Seedance reference/omni and WAN 3.0 reference. WAN 3.0 (wan30-reference): up to 5 clips, 15 seconds combined, each at least 16 fps, addressed in the prompt as "Video 1", "Video 2"; the audio inside the clips is not kept. Seedance: reference video clips for motion/style guidance — up to 10 on Seedance 2.5 (clips 2-30s, 30s combined), up to 3 on Seedance 2.0 (see list_video_models for each model\'s caps). A Seedance video ref can satisfy the required visual anchor. Per the BytePlus guide a video reference is a subject reference: it can carry appearance, identity, motion AND voice timbre. State in the prompt what each asset provides. MiniMax H3 Max reference (h3max-reference): up to 3 clips, each second of reference clip is billed on top of the output seconds (see its notes). Switch Cinema (switch-cinema-480p / switch-cinema-720p): exactly one clip, the one whose motion is kept; the result is as long as this clip and its seconds are the whole price.'}, 'character_orientation': {'enum': ['image', 'video'], 'type': 'string', 'description': 'Motion mode only: follow the character image (default) or the reference video.'}}}
输入模式
{'type': 'object', 'properties': {}}
输入模式
{'type': 'object', 'required': ['layer_set_id'], 'properties': {'layer_set_id': {'type': 'string', 'description': 'The layers project id.'}}}
输入模式
{'type': 'object', 'required': ['layer_set_id'], 'properties': {'layer_set_id': {'type': 'string', 'description': 'The layers project id.'}}}
输入模式
{'type': 'object', 'required': ['text'], 'properties': {'text': {'type': 'string', 'description': 'The words to speak / narrate / perform. Max 3000 chars. For dialogue, address voices as @Audio1, @Audio2, @Audio3. Optional per-sentence timing: put [start:end] seconds in front of a sentence, e.g. [5.5s:8.0s].'}, 'pitch': {'type': 'number', 'description': 'Optional. Pitch, -12 to 12. 0 is normal.'}, 'voice': {'type': 'string', 'description': 'Optional. A saved voice — pass its NAME (or id); it is resolved and routed by kind automatically. Omit for a natural default voice.'}, 'format': {'enum': ['mp3', 'wav', 'ogg_opus'], 'type': 'string', 'description': 'Optional output format. Default mp3.'}, 'delivery': {'type': 'string', 'description': 'Optional. How it should SOUND, in plain English: the voice (gender, age, accent, emotion, tone, speed), mic and room, background sound or music — e.g. "young woman, warm Latin accent, soft, a little flirty, unhurried, clear close mic, no echo". Under 400 chars. Not the words themselves.'}, 'loudness': {'type': 'number', 'description': 'Optional. Loudness, -50 (quieter) to 100 (louder). 0 is normal.'}, 'image_url': {'type': 'string', 'description': 'Optional. Voice a scene from a picture. Cannot be combined with reference audio.'}, 'timestamps': {'type': 'boolean', 'description': "Optional. true returns sentences: [{start, end, text, words: [{start, end, text}]}] in seconds. For a two-voice scene the on-camera person's lines become speak_spans for sync3_lip_sync_video: when you wrote a time window like [3.6s:6.6s] in front of each line, use THOSE windows for her lines (the provider keeps them and they are exact); use sentences and words only for lines you did not time, and check every span holds only her words, because the provider can glue short lines into one sentence."}, 'speech_rate': {'type': 'number', 'description': 'Optional. Speaking speed, -50 (slower) to 100 (faster). 0 is normal.'}, 'reference_audio_url': {'type': 'string', 'description': 'Optional. A short clip URL to instantly match that voice.'}, 'reference_audio_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Optional. Up to 3 reference clip URLs for multi-voice dialogue.'}}}
输入模式
{'type': 'object', 'required': ['subject'], 'properties': {'count': {'type': 'integer', 'maximum': 50, 'minimum': 1, 'description': 'How many images to generate. Default 4. <= 8 returns inline, > 8 queues to Studio. Beta limit: max 50 per request — larger asks are capped at 50 and the response says so.'}, 'model': {'enum': ['GPT Image 2.5', 'Switch Pro', 'Switch Ultra', 'Nano Banana 2', 'Switch Vintage'], 'type': 'string', 'description': 'Optional explicit model: GPT Image 2.5, Switch Pro, Switch Ultra, Nano Banana 2 or Switch Vintage. Nothing else is accepted. Switch Vintage works from a character taught beforehand: pass no reference_image_urls or face_reference_ids with it. If omitted, auto-routed to GPT Image 2.5, Switch Pro or Switch Ultra by subject (see tool description).'}, 'style': {'enum': ['iphone_realism/digital_phone', 'iphone_realism/film_pointshoot', 'iphone_realism/off_duty_intimate', 'movie_scene/neon_noir_action', 'movie_scene/glamour_finance_excess', 'movie_scene/superhero_blockbuster', 'movie_scene/video_game_character', 'movie_scene/generic_action_thriller', 'high_fashion_editorial/testino_glossy', 'high_fashion_editorial/klein_dark_cinematic', 'high_fashion_editorial/inez_vinoodh_hard_flash', 'high_fashion_editorial/leibovitz_painterly', 'high_fashion_editorial/walker_dreamlike', 'high_fashion_editorial/lindbergh_bw_natural', 'high_fashion_editorial/cass_bird_off_duty', 'graphic_editorial_portrait/classic', 'graphic_editorial_portrait/golden_hour', 'travel/hotel_hero', 'travel/rural_property', 'travel/scenic_view', 'travel/drone_aerial', 'travel/lifestyle', 'travel/interior', 'wellness/warm_amber_tropical', 'wellness/hanalei_cinematic', 'wellness/high_key_cyan_beach', 'cinematic_anamorphic/warm_golden', 'cinematic_anamorphic/cool_noir', 'cinematic_anamorphic/moody_desaturated', 'magic_hour_portrait/canon_85mm', 'magic_hour_portrait/hasselblad_80mm', 'product/clean_studio', 'product/lifestyle', 'product/macro_detail', 'product/flat_lay', 'ugc/phone_shot', 'ugc/film_pointshoot', 'ugc/mirror_selfie', 'ugc/car_selfie'], 'type': 'string', 'description': 'Optional curated style stack from the apply_* skill tools. Format "<skill>/<style_key>", e.g. "wellness/warm_amber_tropical" or "high_fashion_editorial/leibovitz_painterly".'}, 'subject': {'type': 'string', 'description': 'Plain-English description of what to generate. E.g. "a woman walking through a hotel lobby" or "morning coffee on the balcony, model wearing a robe". Name the creator\'s staged Studio references as @Image1, @Image2 (their image strip, see get_my_active_references) and they attach automatically; nothing from the strip attaches unless named.'}, 'folder_name': {'type': 'string', 'description': 'Optional Switch Studio folder name. Auto-created if missing. Defaults to the chat-derived title.'}, 'aspect_ratio': {'enum': ['9:16', '16:9', '1:1', '4:5', '5:4', '3:2', '2:3', '4:3'], 'type': 'string', 'description': 'Image aspect ratio. Default 9:16 (vertical, social-friendly). Switch Vintage takes 9:16, 16:9, 4:3, 3:4, 1:1, 2:3 and 3:2 as is; any other ratio is mapped to the nearest one it offers.'}, 'real_photo_look': {'type': 'boolean', 'description': 'Optional. Adds the casual real-photo texture (film grain, amateur iPhone feel). OFF by default — only set true when the user asks for the realistic, unpolished look.'}, 'face_reference_ids': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Face reference asset ids from upload_reference_asset (frame_type "face"). The ONLY way to use a face/likeness reference. Each id is verified server-side (your own untouched original + identity verification) before anything generates or is charged; a URL or generic upload here is rejected.'}, 'reference_image_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': "Optional public image URLs used as GENERIC references (products, scenery, outfits, style). These are never treated as face references — for a person's face/likeness use face_reference_ids."}}}
输出模式
{'type': 'object', 'properties': {'asset': {'type': 'object', 'additionalProperties': True}, 'images': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': True}}, '_widget': {'type': 'object', 'additionalProperties': True}}, 'additionalProperties': True}
输入模式
{'type': 'object', 'required': ['kind', 'description'], 'properties': {'kind': {'enum': ['bed', 'hit'], 'type': 'string', 'description': '"bed" for a music bed under the piece, "hit" for one short sound on a cut.'}, 'format': {'enum': ['mp3', 'wav'], 'type': 'string', 'description': 'Optional output format. Default mp3.'}, 'seconds': {'type': 'number', 'description': 'Bed length in seconds, 4 to 60. Default 15. Ignored for a hit.'}, 'description': {'type': 'string', 'description': 'The feel in plain words, under 400 characters. For example "warm lo-fi beat, 90 bpm, soft kick, no melody" or "short airy whoosh with a soft thump".'}}}
输入模式
{'type': 'object', 'properties': {'mode': {'enum': ['text-to-video', 'image-to-video', 'reference-to-video', 'frame-to-frame', 'motion', 'omni', 'video-edit', 'upscale'], 'type': 'string', 'description': 'Video mode. Must be supported by the chosen model (see list_video_models).'}, 'audio': {'type': 'boolean', 'description': 'Generate audio. ON by default on Seedance 2.5 (text, image and reference), on Omni, and on every WAN 3.0 lane (text, image, reference); set false for a silent clip. Models without audio ignore this. See list_video_models for which models generate audio and the max seconds with vs without audio.'}, 'model': {'type': 'string', 'description': "Model id from list_video_models (e.g. kling-v3, seedance-2.0-t2v, wan-3.0-t2v, h3-max-t2v, topaz). Or prefer option_id from list_video_models. MiniMax H3 Max ids: h3-max-t2v, h3-max-i2v, h3-max-r2v (option ids h3max-text / h3max-image / h3max-reference). MiniMax H3 Max Turbo (preview, about twice as fast at half the price, no reference lane): h3-max-turbo-t2v, h3-max-turbo-i2v (option ids h3max-turbo-text / h3max-turbo-image). WAN 3.0 ids: wan-3.0-t2v, wan-3.0-i2v, wan-3.0-r2v (option ids wan30-text / wan30-image / wan30-reference): 720p or 1080p, any whole second 2 to 30, generated audio on by default, image lane takes image_url and an optional end_image_url, reference lane takes up to 10 reference_image_urls and 5 reference_video_urls. Switch Cinema (option ids switch-cinema-480p / switch-cinema-720p, model switch-cinema-motion): remakes the clip in reference_video_urls with the person in 1 to 8 reference_image_urls, keeps the clip's motion, no duration setting, priced per second of the clip you attach (8 tokens at 480p, 17 at 720p)."}, 'shots': {'type': 'array', 'items': {'type': 'object', 'properties': {'mode': {'enum': ['text-to-video', 'image-to-video', 'reference-to-video', 'frame-to-frame', 'motion', 'omni', 'video-edit', 'upscale'], 'type': 'string', 'description': 'Video mode. Must be supported by the chosen model (see list_video_models).'}, 'audio': {'type': 'boolean', 'description': 'Generate audio. ON by default on Seedance 2.5 (text, image and reference), on Omni, and on every WAN 3.0 lane (text, image, reference); set false for a silent clip. Models without audio ignore this. See list_video_models for which models generate audio and the max seconds with vs without audio.'}, 'model': {'type': 'string', 'description': "Model id from list_video_models (e.g. kling-v3, seedance-2.0-t2v, wan-3.0-t2v, h3-max-t2v, topaz). Or prefer option_id from list_video_models. MiniMax H3 Max ids: h3-max-t2v, h3-max-i2v, h3-max-r2v (option ids h3max-text / h3max-image / h3max-reference). MiniMax H3 Max Turbo (preview, about twice as fast at half the price, no reference lane): h3-max-turbo-t2v, h3-max-turbo-i2v (option ids h3max-turbo-text / h3max-turbo-image). WAN 3.0 ids: wan-3.0-t2v, wan-3.0-i2v, wan-3.0-r2v (option ids wan30-text / wan30-image / wan30-reference): 720p or 1080p, any whole second 2 to 30, generated audio on by default, image lane takes image_url and an optional end_image_url, reference lane takes up to 10 reference_image_urls and 5 reference_video_urls. Switch Cinema (option ids switch-cinema-480p / switch-cinema-720p, model switch-cinema-motion): remakes the clip in reference_video_urls with the person in 1 to 8 reference_image_urls, keeps the clip's motion, no duration setting, priced per second of the clip you attach (8 tokens at 480p, 17 at 720p)."}, 'subject': {'type': 'string', 'description': 'The shot: subject + motion + scene (video needs motion language, e.g. "slow push-in"). Name the creator\'s staged Studio references as @Image1, @Video1 (the Video tab strip, then the image strip; see get_my_active_references) and they attach automatically; nothing attaches unless named. WAN 3.0: refer to references by their place ("Image 1", "Video 1"), not by @tags; there is no negative-prompt field, so write exclusions inline as "Avoid: ...". MiniMax H3 Max: put spoken words in double quotes and the character says them lip-synced; name references by place (Image 1, Video 1, Audio 1).'}, 'duration': {'type': 'string', 'description': 'Clip length in seconds, or "auto" to let the model choose. Seedance 2.5 does 4-30s (auto bills the 30s cap up front and refunds the unused seconds); Seedance 2.0 does 4-15s; Switch Video (WAN 2.7) does 5/10/15; WAN 3.0 takes any whole second from 2 to 30; Kling/Switch Video Edit cap at 10; MiniMax H3 Max takes whole seconds 5 to 15 (no auto) — see each model\'s durations in list_video_models.'}, 'image_url': {'type': 'string', 'description': 'Required for image-to-video / frame-to-frame / motion. Accepts EITHER a Switch asset id (from show_media / list_my_assets / upload_media) OR a public https url. An asset id is resolved server-side, so just pass the id you have — no need to fetch a url first.'}, 'option_id': {'type': 'string', 'description': 'Catalog id from list_video_models. Native Seedance Draft: seedance-25-draft-text, seedance-25-draft-image, seedance-25-draft-reference (480p with audio). Make Full HD: seedance-25-draft-final with the completed draft job_id in video_url, no new creative settings. Quote then generate within 7 days.'}, 'task_type': {'enum': ['auto', 'edit', 'extend'], 'type': 'string', 'description': "Seedance 2.5 reference mode only: declare what you are doing with the reference clip so the size and length rules are checked immediately instead of failing a minute in. auto = a new take from the references (default), edit = change something inside the clip, extend = continue the clip. Editing and extension inherit the source clip's size, and editing also inherits its length."}, 'video_url': {'type': 'string', 'description': 'Required for video-edit and upscale (the source clip). Accepts one of YOUR Switch videos — a job id from list_my_videos / get_video_status, or its download_url / view_url — or any publicly downloadable https URL. Switch resolves its own videos for you; no need to scrape a page for the file.'}, 'resolution': {'enum': ['480p', '720p', '768p', '1080p', '4k'], 'type': 'string', 'description': 'Output resolution. Defaults to 1080p where the model supports it. 720p is cheaper and faster. 480p is the cheapest. Seedance 2.5 offers 480p, 720p and 1080p on every mode (no 2K/4K). 4K is only on Kling v3 text/image and Kling Omni; Seedance text-to-video is 720p only. MiniMax H3 Max offers 480p and 768p only (its reference lane is 768p only). Each model lists its available resolutions in list_video_models.'}, 'aspect_ratio': {'type': 'string', 'description': 'e.g. 9:16, 16:9, 1:1. Must be allowed for the model (see list_video_models).'}, 'end_image_url': {'type': 'string', 'description': 'End frame for frame-to-frame mode.'}, 'quoted_credits': {'type': 'number', 'description': 'Exact quoted_credits returned by quote_video; used with quote_fingerprint.'}, 'quote_fingerprint': {'type': 'string', 'description': 'From quote_video. Pass with quoted_credits to reject a changed quote before charging.'}, 'face_reference_ids': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Face reference asset ids from upload_reference_asset (frame_type "face") — the ONLY way to use a face/likeness reference in video. Each id is verified server-side (your own untouched original + identity verification) before the shot fires or is charged; URLs and generic uploads here are rejected.'}, 'enhancement_receipt': {'type': 'string', 'description': 'Receipt from enhance_video_prompt. Pass its enhanced_prompt as subject with the same model, settings and references. Prevents duplicate enhancement. Changed/expired receipts fail before generation.'}, 'reference_audio_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Seedance reference/omni only: up to 3 reference audio files to drive synthesized audio. Requires at least one reference image or video. MiniMax H3 Max reference (h3max-reference) also takes up to 3 audio files as conditioning (voice/sound guidance); it generates native audio on every clip regardless.'}, 'reference_image_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'GENERIC reference images (products, scenery, outfits, style). Each entry accepts EITHER a Switch asset id (from show_media / list_my_assets / upload_media / get_my_active_references) OR a public https url — asset ids are resolved server-side. Seedance reference/omni accepts up to 9; Kling Omni up to 7. WAN 3.0 reference (option wan30-reference, model wan-3.0-r2v) accepts up to 10; name each one by its place in the prompt ("the woman in Image 1", "the jacket from Image 3"). For Seedance, at least one image or video reference is required. For a person\'s face/likeness use face_reference_ids instead. Switch Cinema (switch-cinema-480p / switch-cinema-720p): 1 to 8 pictures of the person to put into the clip, required.'}, 'reference_video_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Seedance reference/omni and WAN 3.0 reference. WAN 3.0 (wan30-reference): up to 5 clips, 15 seconds combined, each at least 16 fps, addressed in the prompt as "Video 1", "Video 2"; the audio inside the clips is not kept. Seedance: reference video clips for motion/style guidance — up to 10 on Seedance 2.5 (clips 2-30s, 30s combined), up to 3 on Seedance 2.0 (see list_video_models for each model\'s caps). A Seedance video ref can satisfy the required visual anchor. Per the BytePlus guide a video reference is a subject reference: it can carry appearance, identity, motion AND voice timbre. State in the prompt what each asset provides. MiniMax H3 Max reference (h3max-reference): up to 3 clips, each second of reference clip is billed on top of the output seconds (see its notes). Switch Cinema (switch-cinema-480p / switch-cinema-720p): exactly one clip, the one whose motion is kept; the result is as long as this clip and its seconds are the whole price.'}, 'character_orientation': {'enum': ['image', 'video'], 'type': 'string', 'description': 'Motion mode only: follow the character image (default) or the reference video.'}}}, 'description': 'A storyboard of 1-10 DISTINCT shots. Each item takes the same fields as a single shot (subject, model, mode, image_url, etc.).'}, 'subject': {'type': 'string', 'description': 'The shot: subject + motion + scene (video needs motion language, e.g. "slow push-in"). Name the creator\'s staged Studio references as @Image1, @Video1 (the Video tab strip, then the image strip; see get_my_active_references) and they attach automatically; nothing attaches unless named. WAN 3.0: refer to references by their place ("Image 1", "Video 1"), not by @tags; there is no negative-prompt field, so write exclusions inline as "Avoid: ...". MiniMax H3 Max: put spoken words in double quotes and the character says them lip-synced; name references by place (Image 1, Video 1, Audio 1).'}, 'duration': {'type': 'string', 'description': 'Clip length in seconds, or "auto" to let the model choose. Seedance 2.5 does 4-30s (auto bills the 30s cap up front and refunds the unused seconds); Seedance 2.0 does 4-15s; Switch Video (WAN 2.7) does 5/10/15; WAN 3.0 takes any whole second from 2 to 30; Kling/Switch Video Edit cap at 10; MiniMax H3 Max takes whole seconds 5 to 15 (no auto) — see each model\'s durations in list_video_models.'}, 'image_url': {'type': 'string', 'description': 'Required for image-to-video / frame-to-frame / motion. Accepts EITHER a Switch asset id (from show_media / list_my_assets / upload_media) OR a public https url. An asset id is resolved server-side, so just pass the id you have — no need to fetch a url first.'}, 'option_id': {'type': 'string', 'description': 'Catalog id from list_video_models. Native Seedance Draft: seedance-25-draft-text, seedance-25-draft-image, seedance-25-draft-reference (480p with audio). Make Full HD: seedance-25-draft-final with the completed draft job_id in video_url, no new creative settings. Quote then generate within 7 days.'}, 'task_type': {'enum': ['auto', 'edit', 'extend'], 'type': 'string', 'description': "Seedance 2.5 reference mode only: declare what you are doing with the reference clip so the size and length rules are checked immediately instead of failing a minute in. auto = a new take from the references (default), edit = change something inside the clip, extend = continue the clip. Editing and extension inherit the source clip's size, and editing also inherits its length."}, 'video_url': {'type': 'string', 'description': 'Required for video-edit and upscale (the source clip). Accepts one of YOUR Switch videos — a job id from list_my_videos / get_video_status, or its download_url / view_url — or any publicly downloadable https URL. Switch resolves its own videos for you; no need to scrape a page for the file.'}, 'resolution': {'enum': ['480p', '720p', '768p', '1080p', '4k'], 'type': 'string', 'description': 'Output resolution. Defaults to 1080p where the model supports it. 720p is cheaper and faster. 480p is the cheapest. Seedance 2.5 offers 480p, 720p and 1080p on every mode (no 2K/4K). 4K is only on Kling v3 text/image and Kling Omni; Seedance text-to-video is 720p only. MiniMax H3 Max offers 480p and 768p only (its reference lane is 768p only). Each model lists its available resolutions in list_video_models.'}, 'aspect_ratio': {'type': 'string', 'description': 'e.g. 9:16, 16:9, 1:1. Must be allowed for the model (see list_video_models).'}, 'end_image_url': {'type': 'string', 'description': 'End frame for frame-to-frame mode.'}, 'quoted_credits': {'type': 'number', 'description': 'Exact quoted_credits returned by quote_video; used with quote_fingerprint.'}, 'quote_fingerprint': {'type': 'string', 'description': 'From quote_video. Pass with quoted_credits to reject a changed quote before charging.'}, 'face_reference_ids': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Face reference asset ids from upload_reference_asset (frame_type "face") — the ONLY way to use a face/likeness reference in video. Each id is verified server-side (your own untouched original + identity verification) before the shot fires or is charged; URLs and generic uploads here are rejected.'}, 'enhancement_receipt': {'type': 'string', 'description': 'Receipt from enhance_video_prompt. Pass its enhanced_prompt as subject with the same model, settings and references. Prevents duplicate enhancement. Changed/expired receipts fail before generation.'}, 'reference_audio_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Seedance reference/omni only: up to 3 reference audio files to drive synthesized audio. Requires at least one reference image or video. MiniMax H3 Max reference (h3max-reference) also takes up to 3 audio files as conditioning (voice/sound guidance); it generates native audio on every clip regardless.'}, 'reference_image_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'GENERIC reference images (products, scenery, outfits, style). Each entry accepts EITHER a Switch asset id (from show_media / list_my_assets / upload_media / get_my_active_references) OR a public https url — asset ids are resolved server-side. Seedance reference/omni accepts up to 9; Kling Omni up to 7. WAN 3.0 reference (option wan30-reference, model wan-3.0-r2v) accepts up to 10; name each one by its place in the prompt ("the woman in Image 1", "the jacket from Image 3"). For Seedance, at least one image or video reference is required. For a person\'s face/likeness use face_reference_ids instead. Switch Cinema (switch-cinema-480p / switch-cinema-720p): 1 to 8 pictures of the person to put into the clip, required.'}, 'reference_video_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Seedance reference/omni and WAN 3.0 reference. WAN 3.0 (wan30-reference): up to 5 clips, 15 seconds combined, each at least 16 fps, addressed in the prompt as "Video 1", "Video 2"; the audio inside the clips is not kept. Seedance: reference video clips for motion/style guidance — up to 10 on Seedance 2.5 (clips 2-30s, 30s combined), up to 3 on Seedance 2.0 (see list_video_models for each model\'s caps). A Seedance video ref can satisfy the required visual anchor. Per the BytePlus guide a video reference is a subject reference: it can carry appearance, identity, motion AND voice timbre. State in the prompt what each asset provides. MiniMax H3 Max reference (h3max-reference): up to 3 clips, each second of reference clip is billed on top of the output seconds (see its notes). Switch Cinema (switch-cinema-480p / switch-cinema-720p): exactly one clip, the one whose motion is kept; the result is as long as this clip and its seconds are the whole price.'}, 'character_orientation': {'enum': ['image', 'video'], 'type': 'string', 'description': 'Motion mode only: follow the character image (default) or the reference video.'}, 'person_rights_confirmed': {'type': 'boolean', 'description': 'Required with video_people_declaration "person": confirms you own or are authorized to use the person\'s likeness in the reference video.'}, 'video_people_declaration': {'enum': ['none', 'person'], 'type': 'string', 'description': 'Required with reference_video_urls: "none" confirms no real person appears; "person" runs the protected pipeline (verified account, your own stored upload, Seedance 2.5/Mini lane, provider asset registration) — also pass person_rights_confirmed: true.'}}}
输入模式
{'type': 'object', 'properties': {'request_id': {'type': 'string', 'description': 'The request_id from a "rendering" generate_audio answer.'}, 'generation_id': {'type': 'string', 'description': 'Or: the generation_id of a finished take (from generate_audio or list_my_audio).'}}}
输入模式
{'type': 'object', 'required': ['task_id'], 'properties': {'task_id': {'type': 'string', 'description': 'The task_id returned by create_depth_map.'}}}
输入模式
{'type': 'object', 'properties': {'project_id': {'type': 'string'}, 'include_payload': {'type': 'boolean'}}}
输入模式
{'type': 'object', 'required': ['run_id'], 'properties': {'run_id': {'type': 'string'}}}
输入模式
{'type': 'object', 'required': ['layer_set_id'], 'properties': {'layer_set_id': {'type': 'string', 'description': 'The layers project id from split_image_into_layers.'}}}
输入模式
{'type': 'object', 'properties': {}}
输入模式
{'type': 'object', 'properties': {'job_id': {'type': 'string', 'description': 'Alternatively, the job_id.'}, 'task_id': {'type': 'string', 'description': 'Task id returned by generate_video.'}}}
输入模式
{'type': 'object', 'required': ['report_id'], 'properties': {'report_id': {'type': 'string', 'description': 'The report id to fetch.'}}}
输入模式
{'type': 'object', 'required': ['action'], 'properties': {'action': {'enum': ['identify-face', 'create', 'status'], 'type': 'string', 'description': 'Which step to run.'}, 'engine': {'enum': ['best', 'kling'], 'type': 'string', 'description': 'create: OPTIONAL. Leave it out — the default is "best", Sync Labs Sync 3, whole-clip and highest quality, needing only video_url + sound_file. Pass "kling" only for the timeline flow where you place the audio yourself in milliseconds.'}, 'face_id': {'type': 'string', 'description': 'create: a face_id from identify-face (one face supported).'}, 'task_id': {'type': 'string', 'description': 'status: the task_id from create.'}, 'audio_id': {'type': 'string', 'description': 'create: alternative to sound_file — an existing audio id.'}, 'video_url': {'type': 'string', 'description': 'identify-face: the source video (MP4/MOV, 2-60s, <=100MB, 720p/1080p). Use a SwitchApp/public URL.'}, 'session_id': {'type': 'string', 'description': 'create: from identify-face.'}, 'sound_file': {'type': 'string', 'description': 'create: base64 data URI of the audio (e.g. data:audio/mpeg;base64,...).'}, 'speech_volume': {'type': 'number', 'description': 'create: how loud the new speech is, as a percent 0-100 (default 100).'}, 'sound_end_time': {'type': 'integer', 'description': 'create: audio end, in MILLISECONDS.'}, 'keep_background': {'type': 'boolean', 'description': "create, Sync 3 lane: keep the clip's own music and room under the new speech (default true); false gives the new speech alone."}, 'sound_start_time': {'type': 'integer', 'description': 'create: audio start, in MILLISECONDS.'}, 'sound_insert_time': {'type': 'integer', 'description': 'create: where in the video to place the audio, in MILLISECONDS.'}, 'original_audio_volume': {'type': 'number', 'description': "create: how loud the clip's own sound stays, as a percent 0-100 (default 100, the clip keeps its sound)."}}}
输入模式
{'type': 'object', 'properties': {'limit': {'type': 'integer', 'description': 'Default 25, max 100.'}, 'search': {'type': 'string', 'description': 'Optional. Part of a project name.'}}}
输入模式
{'type': 'object', 'properties': {'id': {'enum': ['proof-open', 'voice-narration', 'screen-record', 'real-tray-assets', 'prompt-breakdown-card', 'word-captions', 'beat-and-hits', 'hard-cut-pacing', 'credit-card', 'one-word-cta', 'long-form-retention-cut'], 'type': 'string', 'description': 'Optional workflow id; returns that one workflow in full.'}}}
输入模式
{'type': 'object', 'properties': {'limit': {'type': 'integer', 'description': 'Default 10. Max 50.'}, 'status': {'description': '"all" for everything, or array like ["pending","running"]. Default: active + recent.'}}}
输入模式
{'type': 'object', 'properties': {'count': {'type': 'integer', 'description': 'Default 20. Max 50.'}}}
输入模式
{'type': 'object', 'properties': {'limit': {'type': 'integer', 'description': 'How many takes, newest first. Default 5, max 20.'}, 'search': {'type': 'string', 'description': 'Optional. Words that appear in the take, e.g. a phrase from the script.'}}}
输入模式
{'type': 'object', 'properties': {}}
输入模式
{'type': 'object', 'properties': {'count': {'type': 'integer', 'description': 'How many to return. Default 10. Max 50.'}, 'status': {'type': 'string', 'description': 'Optional filter: submitted, processing, succeed, failed, or all.'}}}
输入模式
{'type': 'object', 'properties': {}}
输入模式
{'type': 'object', 'properties': {}}
输入模式
{'type': 'object', 'properties': {}}
输入模式
{'type': 'object', 'properties': {}}
输入模式
{'type': 'object', 'required': ['id'], 'properties': {'id': {'type': 'string', 'description': 'Playbook id, for example playbook/reference-discipline.'}}}
输入模式
{'type': 'object', 'properties': {'mode': {'enum': ['text-to-video', 'image-to-video', 'reference-to-video', 'frame-to-frame', 'motion', 'omni', 'video-edit', 'upscale'], 'type': 'string', 'description': 'Video mode. Must be supported by the chosen model (see list_video_models).'}, 'audio': {'type': 'boolean', 'description': 'Generate audio. ON by default on Seedance 2.5 (text, image and reference), on Omni, and on every WAN 3.0 lane (text, image, reference); set false for a silent clip. Models without audio ignore this. See list_video_models for which models generate audio and the max seconds with vs without audio.'}, 'model': {'type': 'string', 'description': "Model id from list_video_models (e.g. kling-v3, seedance-2.0-t2v, wan-3.0-t2v, h3-max-t2v, topaz). Or prefer option_id from list_video_models. MiniMax H3 Max ids: h3-max-t2v, h3-max-i2v, h3-max-r2v (option ids h3max-text / h3max-image / h3max-reference). MiniMax H3 Max Turbo (preview, about twice as fast at half the price, no reference lane): h3-max-turbo-t2v, h3-max-turbo-i2v (option ids h3max-turbo-text / h3max-turbo-image). WAN 3.0 ids: wan-3.0-t2v, wan-3.0-i2v, wan-3.0-r2v (option ids wan30-text / wan30-image / wan30-reference): 720p or 1080p, any whole second 2 to 30, generated audio on by default, image lane takes image_url and an optional end_image_url, reference lane takes up to 10 reference_image_urls and 5 reference_video_urls. Switch Cinema (option ids switch-cinema-480p / switch-cinema-720p, model switch-cinema-motion): remakes the clip in reference_video_urls with the person in 1 to 8 reference_image_urls, keeps the clip's motion, no duration setting, priced per second of the clip you attach (8 tokens at 480p, 17 at 720p)."}, 'subject': {'type': 'string', 'description': 'The shot: subject + motion + scene (video needs motion language, e.g. "slow push-in"). Name the creator\'s staged Studio references as @Image1, @Video1 (the Video tab strip, then the image strip; see get_my_active_references) and they attach automatically; nothing attaches unless named. WAN 3.0: refer to references by their place ("Image 1", "Video 1"), not by @tags; there is no negative-prompt field, so write exclusions inline as "Avoid: ...". MiniMax H3 Max: put spoken words in double quotes and the character says them lip-synced; name references by place (Image 1, Video 1, Audio 1).'}, 'duration': {'type': 'string', 'description': 'Clip length in seconds, or "auto" to let the model choose. Seedance 2.5 does 4-30s (auto bills the 30s cap up front and refunds the unused seconds); Seedance 2.0 does 4-15s; Switch Video (WAN 2.7) does 5/10/15; WAN 3.0 takes any whole second from 2 to 30; Kling/Switch Video Edit cap at 10; MiniMax H3 Max takes whole seconds 5 to 15 (no auto) — see each model\'s durations in list_video_models.'}, 'image_url': {'type': 'string', 'description': 'Required for image-to-video / frame-to-frame / motion. Accepts EITHER a Switch asset id (from show_media / list_my_assets / upload_media) OR a public https url. An asset id is resolved server-side, so just pass the id you have — no need to fetch a url first.'}, 'option_id': {'type': 'string', 'description': 'Catalog id from list_video_models. Native Seedance Draft: seedance-25-draft-text, seedance-25-draft-image, seedance-25-draft-reference (480p with audio). Make Full HD: seedance-25-draft-final with the completed draft job_id in video_url, no new creative settings. Quote then generate within 7 days.'}, 'task_type': {'enum': ['auto', 'edit', 'extend'], 'type': 'string', 'description': "Seedance 2.5 reference mode only: declare what you are doing with the reference clip so the size and length rules are checked immediately instead of failing a minute in. auto = a new take from the references (default), edit = change something inside the clip, extend = continue the clip. Editing and extension inherit the source clip's size, and editing also inherits its length."}, 'video_url': {'type': 'string', 'description': 'Required for video-edit and upscale (the source clip). Accepts one of YOUR Switch videos — a job id from list_my_videos / get_video_status, or its download_url / view_url — or any publicly downloadable https URL. Switch resolves its own videos for you; no need to scrape a page for the file.'}, 'resolution': {'enum': ['480p', '720p', '768p', '1080p', '4k'], 'type': 'string', 'description': 'Output resolution. Defaults to 1080p where the model supports it. 720p is cheaper and faster. 480p is the cheapest. Seedance 2.5 offers 480p, 720p and 1080p on every mode (no 2K/4K). 4K is only on Kling v3 text/image and Kling Omni; Seedance text-to-video is 720p only. MiniMax H3 Max offers 480p and 768p only (its reference lane is 768p only). Each model lists its available resolutions in list_video_models.'}, 'aspect_ratio': {'type': 'string', 'description': 'e.g. 9:16, 16:9, 1:1. Must be allowed for the model (see list_video_models).'}, 'end_image_url': {'type': 'string', 'description': 'End frame for frame-to-frame mode.'}, 'quoted_credits': {'type': 'number', 'description': 'Exact quoted_credits returned by quote_video; used with quote_fingerprint.'}, 'quote_fingerprint': {'type': 'string', 'description': 'From quote_video. Pass with quoted_credits to reject a changed quote before charging.'}, 'face_reference_ids': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Face reference asset ids from upload_reference_asset (frame_type "face") — the ONLY way to use a face/likeness reference in video. Each id is verified server-side (your own untouched original + identity verification) before the shot fires or is charged; URLs and generic uploads here are rejected.'}, 'enhancement_receipt': {'type': 'string', 'description': 'Receipt from enhance_video_prompt. Pass its enhanced_prompt as subject with the same model, settings and references. Prevents duplicate enhancement. Changed/expired receipts fail before generation.'}, 'reference_audio_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Seedance reference/omni only: up to 3 reference audio files to drive synthesized audio. Requires at least one reference image or video. MiniMax H3 Max reference (h3max-reference) also takes up to 3 audio files as conditioning (voice/sound guidance); it generates native audio on every clip regardless.'}, 'reference_image_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'GENERIC reference images (products, scenery, outfits, style). Each entry accepts EITHER a Switch asset id (from show_media / list_my_assets / upload_media / get_my_active_references) OR a public https url — asset ids are resolved server-side. Seedance reference/omni accepts up to 9; Kling Omni up to 7. WAN 3.0 reference (option wan30-reference, model wan-3.0-r2v) accepts up to 10; name each one by its place in the prompt ("the woman in Image 1", "the jacket from Image 3"). For Seedance, at least one image or video reference is required. For a person\'s face/likeness use face_reference_ids instead. Switch Cinema (switch-cinema-480p / switch-cinema-720p): 1 to 8 pictures of the person to put into the clip, required.'}, 'reference_video_urls': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Seedance reference/omni and WAN 3.0 reference. WAN 3.0 (wan30-reference): up to 5 clips, 15 seconds combined, each at least 16 fps, addressed in the prompt as "Video 1", "Video 2"; the audio inside the clips is not kept. Seedance: reference video clips for motion/style guidance — up to 10 on Seedance 2.5 (clips 2-30s, 30s combined), up to 3 on Seedance 2.0 (see list_video_models for each model\'s caps). A Seedance video ref can satisfy the required visual anchor. Per the BytePlus guide a video reference is a subject reference: it can carry appearance, identity, motion AND voice timbre. State in the prompt what each asset provides. MiniMax H3 Max reference (h3max-reference): up to 3 clips, each second of reference clip is billed on top of the output seconds (see its notes). Switch Cinema (switch-cinema-480p / switch-cinema-720p): exactly one clip, the one whose motion is kept; the result is as long as this clip and its seconds are the whole price.'}, 'character_orientation': {'enum': ['image', 'video'], 'type': 'string', 'description': 'Motion mode only: follow the character image (default) or the reference video.'}}}
输入模式
{'type': 'object', 'required': ['project_id'], 'properties': {'project_id': {'type': 'string'}}}
输入模式
{'type': 'object', 'required': ['project_id'], 'properties': {'prompt': {'type': 'string', 'description': 'What to build or change, in plain words. Optional when workflow is set.'}, 'workflow': {'enum': ['proof-open', 'voice-narration', 'screen-record', 'real-tray-assets', 'prompt-breakdown-card', 'word-captions', 'beat-and-hits', 'hard-cut-pacing', 'credit-card', 'one-word-cta', 'long-form-retention-cut'], 'type': 'string', 'description': 'Optional named editor workflow to run; its recipe is added to the prompt.'}, 'project_id': {'type': 'string'}}}
输入模式
{'type': 'object', 'required': ['query'], 'properties': {'limit': {'type': 'integer', 'description': 'Default 20.'}, 'query': {'type': 'string'}, 'folderId': {'type': 'string'}}}
输入模式
{'type': 'object', 'required': ['project_id'], 'properties': {'email': {'type': 'string', 'description': "The team member's Switch account email. Required for share and unshare."}, 'action': {'enum': ['share', 'unshare', 'list'], 'type': 'string', 'description': 'Default share.'}, 'project_id': {'type': 'string'}}}
输入模式
{'type': 'object', 'required': ['taskId'], 'properties': {'taskId': {'type': 'string', 'description': 'Task id from generate_image or list_generations.'}}}
输入模式
{'type': 'object', 'properties': {'count': {'type': 'integer', 'description': 'Optional. How many of the most recent images to show as a grid (default 1, max 12). Use when the user says "my last N images/pictures".'}, 'assetId': {'type': 'string', 'description': 'Optional. A specific image id (from list_my_assets, search_my_library, or show_generation). Omit to show the most recent image(s).'}}}
输出模式
{'type': 'object', 'properties': {'asset': {'type': 'object', 'additionalProperties': True}, 'images': {'type': 'array', 'items': {'type': 'object', 'additionalProperties': True}}, '_widget': {'type': 'object', 'additionalProperties': True}}, 'additionalProperties': True}
输入模式
{'type': 'object', 'required': ['asset_id'], 'properties': {'size': {'enum': ['auto', '1K', '1.5K', '2K'], 'type': 'string', 'description': 'Output size for the layers. Auto follows the original image.'}, 'asset_id': {'type': 'string', 'description': 'The id of your image to separate. "my last image" also works.'}, 'instruction': {'type': 'string', 'description': 'Optional: what to separate, e.g. "the outfit and the bunny". Leave empty to separate everything the model finds.'}}}
输入模式
{'type': 'object', 'required': ['clip_asset_ids'], 'properties': {'fps': {'type': 'integer', 'description': 'Frames per second. Default 30.'}, 'quality': {'type': 'string', 'description': 'draft, standard (default), or high.'}, 'orientation': {'type': 'string', 'description': 'landscape (1920x1080, default), portrait (1080x1920), or square (1080x1080).'}, 'project_name': {'type': 'string', 'description': 'Optional name for the output video.'}, 'clip_asset_ids': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Ordered list of your video ids (from list_my_videos). At least 2. Output order = this order.'}}}
输入模式
{'type': 'object', 'properties': {'text': {'type': 'string', 'description': 'The spoken words. The tool counts them for you.'}, 'word_count': {'type': 'integer', 'description': 'Use instead of text if you already know the number of spoken words.'}}}
输入模式
{'type': 'object', 'required': ['action'], 'properties': {'action': {'enum': ['create', 'status'], 'type': 'string', 'description': 'Create a Sync 3 lip-sync job or check its status.'}, 'task_id': {'type': 'string', 'description': 'status: the sync: task_id returned by create.'}, 'audio_url': {'type': 'string', 'description': 'create: public or Switch-resolvable URL of the audio. Use instead of sound_file.'}, 'video_url': {'type': 'string', 'description': 'create: source video URL.'}, 'sound_file': {'type': 'string', 'description': 'create: base64 data URI of the audio (for example data:audio/mpeg;base64,...).'}, 'speak_spans': {'type': 'array', 'items': {'type': 'object', 'required': ['start', 'end'], 'properties': {'end': {'type': 'number'}, 'start': {'type': 'number'}}}, 'description': "create, two-voice scenes: the seconds when the person ON CAMERA speaks: the time windows you wrote in front of her lines in generate_audio's text (exact), or from its sentences and words (timestamps: true) when a line was not timed. Switch silences everything else in a copy of the audio so the mouth follows only those lines; the other voice never moves it. Up to 60 spans, in order, no overlap."}, 'keep_background': {'type': 'boolean', 'description': "create: keep the clip's own music and room under the new speech (default true). Switch splits the soundtrack into layers and lays the non-voice layers back; only the voice changes. false gives the new speech alone."}, 'scene_audio_url': {'type': 'string', 'description': 'create: the full scene mix (both voices and the room) to lay onto the synced picture when it is done, usually the same audio_url. With speak_spans this makes one dialogue do both jobs: drive the lips and play underneath. Replaces keep_background.'}}}
输入模式
{'type': 'object', 'required': ['image_url'], 'properties': {'script': {'type': 'string', 'description': 'Text the avatar speaks. Max 2500 characters. Required unless audio_url is given.'}, 'language': {'type': 'string', 'description': 'Optional language code (default en).'}, 'voice_id': {'type': 'string', 'description': 'Optional voice id (from clone_voice / your library).'}, 'audio_url': {'type': 'string', 'description': 'Pre-recorded audio URL to lip-sync instead of generating speech from script.'}, 'image_url': {'type': 'string', 'description': 'A clear face photo (Switch/public URL). Required.'}, 'voice_settings': {'type': 'object', 'description': 'Optional: { stability, similarityBoost, style, useSpeakerBoost } 0-1.'}}}
输入模式
{'type': 'object', 'required': ['layer_set_id', 'changes'], 'properties': {'changes': {'type': 'array', 'items': {'type': 'object', 'required': ['layer_id'], 'properties': {'box': {'type': 'array', 'items': {'type': 'number'}, 'description': 'Move and resize: [left, top, right, bottom] in base-image pixels.'}, 'order': {'type': 'integer', 'description': 'Stacking order; higher sits on top. The base always stays at the bottom.'}, 'reset': {'type': 'boolean', 'description': 'Put this layer back exactly where the split placed it.'}, 'visible': {'type': 'boolean', 'description': 'Show or hide this layer.'}, 'layer_id': {'type': 'string', 'description': 'The layer id from get_layer_set.'}}}, 'description': 'One entry per layer you are changing.'}, 'layer_set_id': {'type': 'string', 'description': 'The layers project id.'}}}
输入模式
{'type': 'object', 'required': ['project_id', 'kind'], 'properties': {'kind': {'enum': ['image', 'video', 'audio'], 'type': 'string'}, 'mime': {'type': 'string'}, 'presign': {'type': 'boolean'}, 'filename': {'type': 'string'}, 'file_size': {'type': 'integer', 'description': 'Bytes; required for presign.'}, 'project_id': {'type': 'string'}, 'confirm_path': {'type': 'string'}, 'duration_seconds': {'type': 'number'}}}
输入模式
{'type': 'object', 'properties': {'url': {'type': 'string', 'description': 'Any public https URL — Switch fetches it server-side.'}, 'mime': {'type': 'string', 'description': 'MIME type when sending base64. Default image/png.'}, 'base64': {'type': 'string', 'description': 'Base64-encoded image bytes (use this when there is no public URL).'}}}
输入模式
{'type': 'object', 'required': ['kind'], 'properties': {'url': {'type': 'string', 'description': 'Public https URL to fetch server-side.'}, 'kind': {'enum': ['image', 'video', 'audio'], 'type': 'string', 'description': 'Reference type to upload.'}, 'mime': {'type': 'string', 'description': 'MIME for base64. Images: jpg/png/webp/gif. Videos: mp4/mov. Audio: mp3/wav/m4a/aac.'}, 'base64': {'type': 'string', 'description': 'Base64 bytes (optionally a data: URL). Best for small files; large video should use presign.'}, 'presign': {'type': 'boolean', 'description': 'Return an upload_url to PUT the file bytes directly to (no base64). Video always; image/audio when enabled.'}, 'activate': {'type': 'boolean', 'default': False, 'description': 'Image only. Set true only for an explicit user request to add this image to the Studio image workspace. Default false: upload and return a reusable reference without changing Studio. On presigned uploads, set this on the confirm_path call. Video and audio never activate Studio.'}, 'filename': {'type': 'string', 'description': 'Optional source filename for extension/display.'}, 'frame_type': {'type': 'string', 'description': 'Image strip label: ref (default), face, body, clothes, scenery, product, typography. Use "face" for a person\'s face/likeness — face uploads are stored as untouched originals in the private reference bucket and their returned asset_id is the ONLY handle face-capable generation accepts (KYC-verified accounts only).'}, 'confirm_path': {'type': 'string', 'description': 'The storage_path from a presign call, after you PUT the file — verifies the object and records it; only activate=true on this call adds an image to Studio.'}}}
输入模式
{'type': 'object', 'properties': {'model': {'enum': ['Proteus', 'Artemis HQ', 'Artemis MQ', 'Artemis LQ', 'Nyx', 'Nyx Fast', 'Nyx XL', 'Nyx HF', 'Gaia HQ', 'Gaia CG', 'Gaia 2', 'Starlight Precise 2.5', 'Starlight HQ', 'Starlight Mini', 'Starlight Sharp', 'Starlight Fast 2', 'Proteus Natural', 'Iris', 'Iris Low Quality', 'Dione DV', 'Dione TV', 'Dione Robust', 'Dione Dehalo', 'Dione Robust Dehalo', 'Artemis Strong Halo', 'Artemis Medium Halo', 'Artemis Aliasing & Moire', 'Rhea', 'Theia Fine Tune Detail', 'Theia Fine Tune Fidelity', 'Starlight Precise 2.6', 'FlashVSR', 'Real-ESRGAN', 'BRIA Increase Resolution', 'Starlight Fast 3'], 'type': 'string', 'description': 'Upscaler model. Default Proteus. Use list_upscalers for current availability and model specific controls. Fast 3 uses the direct Topaz API; Precise 2.5 remains available.'}, 'scale': {'type': 'number', 'description': 'Upscale factor, 1 to 4. 1.5 is allowed. Default 2.'}, 'video': {'type': 'string', 'description': 'One clip: a Switch job id, its download_url/view_url, or a public https video URL.'}, 'videos': {'type': 'array', 'items': {'type': 'string'}, 'description': 'Batch of up to 10 clips (same accepted forms as video).'}, 'quote_only': {'type': 'boolean', 'description': 'Return exact token quotes without starting processing or deducting tokens. Requires a saved Switch clip.'}, 'request_id': {'type': 'string', 'description': 'Stable identifier for one intended batch. Reuse across retries and reloads; use a new identifier only for a deliberate new run.'}, 'target_fps': {'type': 'integer', 'description': 'Optional target frame rate 16 to 60 (enables frame interpolation; 60 doubles the token cost).'}, 'h264_output': {'type': 'boolean', 'description': 'Use H.264 for broad playback compatibility. Defaults false, preserving the existing H.265 option.'}, 'confirmed_quotes': {'type': 'array', 'items': {'type': 'object', 'required': ['tokens', 'fingerprint'], 'properties': {'tokens': {'type': 'number'}, 'fingerprint': {'type': 'string'}}}, 'description': 'Exact quote receipts in the same order as videos.'}}}
输入模式
{'type': 'object', 'properties': {'mode': {'enum': ['action', 'scene', 'action_scene', 'description_scene'], 'type': 'string', 'description': "What the prompt focuses on. action = motion/gestures only, no appearance or scene. scene = setting/camera/lighting only, no subject. action_scene = both, no appearance. description_scene (default) = full prompt including the subject's appearance."}, 'engine': {'enum': ['seedance', 'kling', 'gemini'], 'type': 'string', 'description': 'Which model format to return: seedance (default, control-format), kling (cinematic prose), or gemini (plain paragraph for Omni).'}, 'report_id': {'type': 'string', 'description': 'A finished report id from analyze_video_report or list_vision_reports.'}, 'video_url': {'type': 'string', 'description': 'Alternative: the exact public https URL you already analyzed.'}}}
输入模式
{'type': 'object', 'required': ['action'], 'properties': {'action': {'enum': ['clone', 'list', 'delete', 'create', 'rename'], 'type': 'string', 'description': 'Which operation to run.'}, 'format': {'enum': ['wav', 'mp3', 'ogg', 'm4a', 'aac'], 'type': 'string', 'description': 'create/clone: clip format when sending audio_base64. Default wav.'}, 'language': {'type': 'string', 'description': 'clone: language of the sample, e.g. en, es, ja. Default en.'}, 'new_name': {'type': 'string', 'description': 'rename: the new name for the voice.'}, 'voice_id': {'type': 'string', 'description': 'delete/rename: the voice id OR its name — names are resolved for you.'}, 'audio_url': {'type': 'string', 'description': 'create: URL of a 10-30 second clip of the voice — e.g. the url returned by upload_media.'}, 'voice_name': {'type': 'string', 'description': 'create/clone: what to call the voice (unique per account).'}, 'audio_base64': {'type': 'string', 'description': 'create/clone: the clip as base64 when there is no URL.'}, 'audio_sample_url': {'type': 'string', 'description': 'clone: URL of a clean 10-15 second clip of one person talking (reachable).'}}}
近期工具变更
类似的 MCP 服务器
MiOffice — AI-Powered Workspace Studio
Provides browser-based tools for processing PDFs, images, video, and audio, including generation, enhancement, conversion, transc…
BlitzReels Video Editor
Creates and edits short-form videos, including timelines, captions, transitions, AI-generated visuals, music, voiceovers, sound e…
Morpha
Enables creating, animating, and exporting layered short-form video projects.
framesail
Creates long-form YouTube videos through script generation, storyboarding, visual assets, voiceover, music, scene composition, an…
Magnific
Provides image enhancement and design tools, AI image and audio generation, text-to-speech, creation management, and reusable cre…
Uwear
Supports AI fashion photoshoot production with garments, outfits, avatars, locations, art direction, image generation, backdrops,…
Creative Claw
Generates and edits branded images, video, audio, speech, and 3D-related creative assets, with media processing and reusable them…
RegiAI
Generates and edits images, videos, audio, music, speech, subtitles, avatars, and related media asynchronously.