Cloud

A vision API for people in images and video.

Hosted AI inference engines, returned as structured JSON: faces, expressions, head and body pose, identities, transcripts.

Start freeRead the docsSign in

$5 of free credit for new accounts. Pay per task after that. Pricing · Live demo

Images

Faces, expressions and pose from a single image.

Send a photo with the tasks you want and get every face back with its box, landmarks, expression scores and head pose, plus body keypoints and identity matches when you ask for them.

terminal
curl https://api.beemotion.app/v1/image/analyze \
  -H "Authorization: Bearer $BEEMOTION_API_KEY" \
  -F [email protected] \
  -F tasks=face,emotion
Video

Video analysis with timelines and transcripts.

Submit a video as an async job and get a signed webhook when it finishes. Faces become tracks, each with its own expression timeline, and the audio track becomes timed transcript segments.

webhook payload
{
  "object": "video.prediction",
  "status": "completed",
  "tasks": ["face", "emotion", "pose", "transcript"],
  "summary": {
    "tracks": [ { "track_id": 1, … } ],
    "transcript": {
      "language": "en",
      "segments": [
        { "start": 0.4, "end": 5.1, "text": "Okay, we'll reset…" }
      ]
    }
  },
  "usage": { "units": …, "cost_usd": … }
}
Agents

An MCP server and an agent skill.

Add the MCP server to Claude Code, Cursor or Codex with your API key and install the agent skill next to it. Agent calls are billed and logged like REST calls.

terminal
claude mcp add --transport http beemotion \
  https://api.beemotion.app/mcp \
  --header "Authorization: Bearer $BEEMOTION_API_KEY"

# Then ask your agent:
# "How many faces are in photo.jpg, and what expression is each one showing?"
Capabilities

Every task the API runs.

Eleven capabilities for people in images and video. Ask for several tasks in one call; each tile links to its page in the docs.

Recognition only matches identities you enrol, with consent. Privacy

Platform

One API for scripts, services and agents.

Use cases

What teams build with it.

A sales conversation across a table with a laptop
Sales coaching

Find the moments a pitch lands

Sales managers review recorded pitches to find where a prospect showed interest or hesitation, with the transcript alongside, and coach to those moments.

A cinema audience laughing during a screening
Media production

Measure audience reaction scene by scene

Studios record test screenings and follow audience expression through each scene, so editors can compare reactions between cuts.

A player wearing a headset at a gaming PC
Gaming

Adapt pacing to the player

Game developers send webcam frames to the image endpoint and use the player's expression and head pose to adjust difficulty and narrative pacing.

Financial market data on a screen
Banking

Train advisers on real consultations

Financial institutions review recorded client consultations to find moments where clients showed confusion or concern, and use them in adviser training.

Travellers at automated passport gates in an airport
Security

Review checkpoint footage faster

Security teams run checkpoint footage through face detection, head pose and object detection to find the segments worth a closer look, instead of watching hours of video.

For pilot training and driver monitoring, see AeroMind and In-Cabin sensing.

BeEmotion models also run on edge devices, for cameras and vehicles where video should stay on site.

Explore Edge →