
Find the moments a pitch lands
Sales managers review recorded pitches to find where a prospect showed interest or hesitation, with the transcript alongside, and coach to those moments.
Send a photo with the tasks you want and get every face back with its box, landmarks, expression scores and head pose, plus body keypoints and identity matches when you ask for them.
curl https://api.beemotion.app/v1/image/analyze \
-H "Authorization: Bearer $BEEMOTION_API_KEY" \
-F [email protected] \
-F tasks=face,emotionSubmit a video as an async job and get a signed webhook when it finishes. Faces become tracks, each with its own expression timeline, and the audio track becomes timed transcript segments.
{
"object": "video.prediction",
"status": "completed",
"tasks": ["face", "emotion", "pose", "transcript"],
"summary": {
"tracks": [ { "track_id": 1, … } ],
"transcript": {
"language": "en",
"segments": [
{ "start": 0.4, "end": 5.1, "text": "Okay, we'll reset…" }
]
}
},
"usage": { "units": …, "cost_usd": … }
}Add the MCP server to Claude Code, Cursor or Codex with your API key and install the agent skill next to it. Agent calls are billed and logged like REST calls.
claude mcp add --transport http beemotion \
https://api.beemotion.app/mcp \
--header "Authorization: Bearer $BEEMOTION_API_KEY"
# Then ask your agent:
# "How many faces are in photo.jpg, and what expression is each one showing?"Eleven capabilities for people in images and video. Ask for several tasks in one call; each tile links to its page in the docs.
Who is in the frame, where they are looking, their expression and how they move.
The objects around them, from a fixed set or from labels you name.
Text in the image, descriptions, transcripts and questions in plain language.
Recognition only matches identities you enrol, with consent. Privacy

Sales managers review recorded pitches to find where a prospect showed interest or hesitation, with the transcript alongside, and coach to those moments.

Studios record test screenings and follow audience expression through each scene, so editors can compare reactions between cuts.

Game developers send webcam frames to the image endpoint and use the player's expression and head pose to adjust difficulty and narrative pacing.

Financial institutions review recorded client consultations to find moments where clients showed confusion or concern, and use them in adviser training.

Security teams run checkpoint footage through face detection, head pose and object detection to find the segments worth a closer look, instead of watching hours of video.
For pilot training and driver monitoring, see AeroMind and In-Cabin sensing.
BeEmotion models also run on edge devices, for cameras and vehicles where video should stay on site.
Explore Edge →