Benchmark methodology

Measured through the public partner API

The benchmark behaves like an external integration: it discovers live capabilities, uploads approved inputs, creates asynchronous jobs, polls public status, verifies output delivery, and optionally observes signed webhook delivery. It does not call private queues, databases, workers, or provider handlers.

Test matrix

  • Image-to-video and reference-guided video only when returned by live capabilities
  • 5, 10, and 15 second clips when each duration is currently available
  • 720p and 1080p when currently available; 4K is excluded from this benchmark policy
  • Smoke: 3, preliminary: 10, and formal: 30 samples per configuration
  • Separate concurrency batches of 1, 3, and 5 simultaneous requests

Measurements

Lifecycle timing is calculated from API acceptance and externally observed public stages: queue wait, generation, finalization, total result latency, and receiver-confirmed webhook latency. Polling interval is reported as bounded observation error. Cold or warm worker state remains unknown unless reliable worker evidence exists.

Publication threshold

Preliminary results require at least 10 submitted samples per tested configuration. Formal per-configuration p95 results require at least 30. Missing, invalid, unvalidated, or undersized reports never produce public numerical claims.

Cost and reproducibility

Paid execution requires explicit authorization plus independent maximum-credit and estimated-infrastructure-cost ceilings. Reports record the service version, deployment revision, matrix, skips, formulas, and declared environment while the public summary excludes credentials, private endpoints, signed URLs, input paths, and provider details.

Partner Benchmark Methodology | MotionLab