Kling video and image API
The strict Tencent VOD contract for Kling video, image, motion-control, avatar, and lip-sync tasks.
Essevin exposes Kling through Tencent Cloud VOD's AIGC gateway. It is an asynchronous API: submit a task, then poll the Essevin task ID. Use a key that has the Kling VOD account enabled and confirm the exact model IDs with GET /v1/models before sending paid work.
Endpoints
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/images/generations | Submit a Kling image or image-expand task |
| POST | /v1/kling/faces | Detect faces before lip sync; billed per call |
| POST / GET | /v1/kling/subjects | Create / list custom subjects |
All requests use Authorization: Bearer sk-your-key and Content-Type: application/json. The body is decoded strictly. Unknown fields, a second JSON value, top-level image_url/video_url, and unverified nested fields are rejected before Tencent is called.
Model matrix
| Model ID | Duration | Reference images | Reference video | Subjects | Shots | Notes |
|---|---|---|---|---|---|---|
kling-v3-turbo | 3-15 s | No | No | Yes | No | Fixed Tencent Voice price bucket; callers cannot select a voice |
kling-v3-omni | 3-15 s | Up to 8 | feature | Yes | Yes | base is not priced; 4K silent feature is not currently priced |
kling-v3 | 3-15 s | Up to 6 | No | Yes | Yes | Ordinary generation and image references |
kling-o1 | 3-10 s | Up to 4 | feature | No | No | With no reference generation input, duration is 5 or 10 s |
kling-v2-6 | 5-10 s | Up to 4 | No | No | No | Audio 720P is not priced |
kling-v2-5-turbo | 5-10 s | Up to 3 | No | No | No | 720P / 1080P only |
kling-v2-1, kling-v2-0 | 5-10 s | No | No | No | No | Ordinary text-to-video only |
kling-v1-6 | 5-10 s | No | base | No | No | Base video editing uses multi_elements pricing |
kling-v3-motion-control | source duration | 1 image + 1 video | Scene-specific | No | No | 720P / 1080P / 2K / 4K |
kling-v2-6-motion-control | source duration | 1 image + 1 video | Scene-specific | No | No | 720P / 1080P |
kling-avatar | source duration | 1-5 images | No | No | No | sound_file XOR audio_id |
kling-lip-sync | source duration | No | No | No | No | session_id + one face_choose item |
Kling image IDs are kling-image-v3, kling-image-v3-omni, kling-image-o1, kling-image-v2-1, kling-image-v2-1-i2i, kling-image-v2-1-multi-ref, and kling-image-expand. They use the image endpoint, accept n from 1 to 9, and are billed per output image. Use the model catalog for the exact quality tiers available to the current key.
Request contract
Media locators
Each images[] or videos[] item must contain exactly one of:
{ "url": "https://cdn.example.com/file.png" }or
{ "file_id": "vod-file-id" }url must be an absolute, publicly reachable http:// or https:// URL. Empty strings, relative paths, file://, ftp://, and both fields together are rejected. The same rule applies to the image endpoint and face detection.
Ordinary video generation
Use images[] for frames or image references and videos[] for one reference/edit video. Do not use OpenAI's image_url or video_url fields at the top level.
{
"model": "kling-v3-omni",
"prompt": "A product rotates slowly on a clean studio table",
"duration": 5,
"resolution": "1080p",
"audio": false,
"images": [
{ "url": "https://cdn.example.com/first.png", "usage": "first_frame" },
{ "file_id": "vod-last-frame", "usage": "last_frame" },
{ "url": "https://cdn.example.com/reference-a.png", "usage": "reference" },
{ "file_id": "vod-reference-b", "usage": "reference" }
],
"videos": [
{
"url": "https://cdn.example.com/character-motion.mp4",
"reference_type": "feature",
"keep_original_sound": false
}
],
"subjects": [
{ "id": "subject-92951593344", "name": "cat" },
{ "id": "subject-92951593345", "name": "dog" }
]
}For ordinary generation, images[].usage is required and is one of first_frame, last_frame, or reference. There can be at most one first frame and one last frame; a last frame requires a first frame. When more than two reference images are present, a last frame is not supported. There is at most one videos[] item, and its reference_type is required. A feature video is accepted only by kling-v3-omni and kling-o1. A base video is accepted only by kling-v1-6; it must be the only media input and cannot be combined with images or subjects.
Tencent's coupling limits also apply: with a reference video, reference image count plus subject count is at most 4; without a reference video it is at most 7. Empty prompts are allowed only when a media input supplies the request. Omitted duration, resolution, and audio normalize to 5 seconds, 720P, and silent output.
Subjects and shots
subjects[] uses fixed Tencent subject IDs. Every item requires a non-empty id; name is optional. Subjects are supported by kling-v3-turbo, kling-v3, and kling-v3-omni only. The prompt references them using the Tencent subject syntax, for example <<<element_1>>>.
Shots are supported by kling-v3 and kling-v3-omni only:
{
"model": "kling-v3-omni",
"prompt": "",
"duration": 5,
"shots": {
"mode": "customize",
"segments": [
{ "index": 1, "prompt": "The box opens", "duration": 2 },
{ "index": 2, "prompt": "The product is revealed", "duration": 3 }
]
}
}mode is intelligence or customize. Intelligence mode must omit segments; customize mode requires them. Custom segments are numbered consecutively from 1, have non-empty prompts of at most 512 characters, last at least one second, and their durations must sum exactly to the request duration. Use structured shots; raw extra.multi_shot, extra.shot_type, and extra.multi_prompt are rejected.
Motion control
Motion-control tasks require exactly one video followed by exactly one person image. The scene determines their meaning, so do not send usage on the image or reference_type on the video. videos[].keep_original_sound is a boolean and is mapped to Tencent's keep_original_sound flag. The only additional parameter currently verified is extra.character_orientation (image or video). Motion-control duration comes from the input video and must be omitted; the output is temporary.
{
"model": "kling-v3-motion-control",
"prompt": "Follow the dancer's movement",
"resolution": "1080p",
"images": [{ "file_id": "vod-person-image" }],
"videos": [{ "url": "https://cdn.example.com/dance.mp4", "keep_original_sound": true }],
"extra": { "character_orientation": "video" }
}Avatar and lip sync
Avatar (kling-avatar) requires 1-5 person images, no videos, and exactly one of extra.sound_file (an HTTP(S) audio URL) or extra.audio_id. The duration is derived from the audio input; omit duration.
Lip sync (kling-lip-sync) accepts no images. First call face detection with one or more media locators:
curl https://api.essevin.com/v1/kling/faces \
-H "Authorization: Bearer sk-your-key" \
-H "Content-Type: application/json" \
-d '{"videos":[{"file_id":"vod-source-video"}]}'Then submit one face_choose item for the returned session:
{
"model": "kling-lip-sync",
"extra": {
"session_id": "face-session-id",
"face_choose": [
{ "face_id": "face-1", "sound_file": "https://cdn.example.com/voice.mp3" }
]
}
}face_choose must contain exactly one item with a non-empty face_id and sound_file. The source audio/video determines duration. Lip-sync billing is per second with a five-second minimum, so a four-second result is billed as five seconds. Raw extra.voice_list and voice_ids are intentionally unavailable until Tencent confirms a matching SKU.
Submit and poll
-H "Authorization: Bearer sk-your-key" \
-H "Content-Type: application/json" \
-d '{"model":"kling-v3","prompt":"A red ball rolls across a white table","duration":5,"resolution":"720p"}'The 202 response contains an ID such as kt57x<upstream-task-id>. Poll it without a billing header:
-H "Authorization: Bearer sk-your-key"Statuses are queued, processing, completed, and failed. Completed outputs are Essevin relay URLs; the source is not stored permanently by the plugin. The relay validates the task's file token before fetching Tencent's URL. Billing is produced once by the internal pinned poller only after a successful Tencent FINISH response with ErrCode=0, a valid output, and a billable duration. FINISH with a non-zero ErrCode or ErrCodeExt is a failure and is never charged.
Pricing units and buckets
The model catalog publishes price.unit: video is second, images are image, and face detection is call. Video baseline prices are Tencent's CNY price divided by the site's fixed 6.8 CNY/USD rate; Core applies the group multiplier separately.
| Request shape | Billing bucket | Unit |
|---|---|---|
| Ordinary video | <resolution>_<silent|audio|voice>_<noref|ref> | USD / second |
| Motion control | motion_control_<resolution> | USD / second |
| Avatar | avatar_<resolution> | USD / second |
| Lip sync | lip_sync | USD / second, 5-second minimum |
kling-v1-6 base edit | multi_elements_<resolution> | USD / second |
| Kling image | img_<1k|2k|4k> | USD / image |
| Face detection | call_face_detect | USD / call |
Missing or unconfirmed buckets are unavailable rather than silently mapped to a nearby price. In particular, kling-v3-omni 4K silent feature-video pricing is withheld pending a second Tencent confirmation. Caller-selected voices are also withheld: do not send voice_ids or extra.voice_list.