Back to Home
Product 03 — AI APIs

AI Ability Tools

A comprehensive suite of 260+ AI APIs covering optical character recognition, speech recognition and synthesis, natural language processing, and computer vision. RESTful APIs with SDKs for every major language — integrate AI in minutes.

Capability Radar 260+ APIS ONLINE
260+
APIS
📄
OCR & Docs
40+
🎤
Speech
30+
💬
NLP
60+
👀
Vision
50+
🔖
Knowledge
30+
Workflow
50+
License Plate Handwriting OCR Voice Clone Sentiment Face Search NER Summarize Object Detect Doc Translate Speech 2 Text Text 2 Speech Entity Link

API Categories

260+ production-ready AI APIs across six domains

F01
OCR & Document
General text recognition, handwriting, table extraction, receipt/invoice parsing, ID card scanning, and structured document understanding with layout analysis.
F02
Speech & Audio
Real-time and batch speech recognition in 20+ languages, neural TTS with custom voice cloning, speaker verification, and audio event detection.
F03
Natural Language
Sentiment analysis, entity recognition, text classification, machine translation (100+ language pairs), text summarization, and semantic search.
F04
Computer Vision
Object detection, facial recognition, body tracking, image classification, scene understanding, and visual search with product and logo recognition.
F05
Knowledge Graph
Entity linking, relation extraction, knowledge base Q&A, and reasoning APIs. Pre-built industry knowledge graphs for finance, healthcare, and legal.
F06
AI Workflow
Drag-and-drop pipeline builder to chain multiple AI APIs. Conditional logic, batch processing, and webhooks for automated business workflows.

01 — Text Recognition (OCR)

Multi-scenario, multi-language, high-accuracy text recognition for document digitization and content moderation — available as offline SDK and cloud API

1.1
OCR SDK
Embed OCR into mobile devices (Windows, Android, iOS) — phones, tablets, cameras, industrial controllers. Works fully offline in weak or no-network environments: underground parking, closed warehouses, factory production lines. Not for server-side deployment (x86 Linux, Windows Server).
OVERSEAS READY
1.2
OCR API
Cloud OCR API — overseas node deployment not yet available; by-case discussion possible. Recommended direction: PaddleOCR — recognizes printed text, handwriting, tables, formulas, charts, and seals across 111 languages (Chinese, English, Japanese, Korean, Latin and more). Infers reading order intelligently, outputs structured element sequences with row-level coordinates. Handles irregular layouts and long cross-page documents.
PLANNED

02 — Face Recognition

On-device face SDK for offline identity verification; cloud face API in planning

2.1
Face Offline SDK
On-device localized face detection and capture, multi-modal liveness detection, face comparison and recognition. Runs fully offline for identity verification, driver status analysis, attention detection, and facial attribute analysis.
OVERSEAS READY
2.2
Face Recognition API
Cloud face API — overseas node deployment not yet available; by-case discussion possible. Priority directions: face search, liveness detection, and related anti-fraud capabilities.
PLANNED

03 — Speech Products

Offline TTS SDK, on-device SDKs, and LLM-powered speech APIs — with proven Europe and North America node deployments

3.1
Offline TTS SDK
Converts text to speech on-device in no-network or weak-network environments — mobile apps, e-readers, robots, and smart hardware terminals. Stable, consistent, and naturally fluent synthesis experience.
OVERSEAS READY
3.2
On-Device SDK / Cloud API
Model deployment on Europe and North America nodes (via overseas vendor resources) is proven and available for new projects after cost evaluation. Includes LLM speech products: end-to-end speech language model, voice cloning, and LLM-powered audio file quality inspection — plus classic ASR and TTS.
OVERSEAS READY

04 — Translation Products

LLM-powered and classic translation APIs for cross-border e-commerce, product globalization, and smart hardware — 200+ languages with custom terminology

4.1
LLM Text Translation API
Translation interface powered by Baidu's proprietary multilingual LLM: quality exceeds machine translation, response speed far exceeds general LLMs, instruction-tuned for stable output, custom terminology support for domain precision.
4.2
Text Translation API
Input text, output translated text. 200+ language pairs, general-purpose and domain-specific interfaces, custom terminology for precision.
4.3
Document Translation API
Input documents, output translated documents. Supports Word, PPT, Excel, HTML, XML, TXT, PDF across 200+ languages — with layout preservation in the translated output.
4.4
Image Translation API
Translate text within images — OCR recognition supports 20 languages, with translated text rendered back into the original image scene.
4.5
Speech Translation API
Input short speech, output translated text and audio. Source-language input supports 10+ languages — speech recognition converts audio to text, translates to the target language, and supports spoken playback of the translation.
KEY
Overseas Delivery Models
Public cloud translation API for cross-border e-commerce and global products; private deployment projects are the current main overseas channel; public cloud overseas rollout is on the product roadmap, with customer POCs already completed. Reach out for project-specific discussion.
OVERSEAS READY

05 — Baidu Input Method (Offline)

AI input method for smart terminals — 104 supported languages, delivered as APK, NRE + per-license pricing

5.1
Watch Input Method
For smart watch and kids watch manufacturers: handwriting input, short-speech recognition, Pinyin, quick emojis, and quick symbols — compact interaction for small screens.
5.2
Automotive Input Method
For vehicle OEMs and infotainment vendors: handwriting and Pinyin input; handwriting supports mixed Chinese-English-numeric input with tilt tolerance; custom geographic and music dictionaries; floating mode support.
5.3
TV Input Method
For TV manufacturers and TV OS vendors: custom movie and celebrity dictionaries, cross-device input between phone and TV for ten-foot interaction.
5.4
Large-Screen Input Method
For conference display manufacturers: dedicated large-screen layout with floating mode for meeting-room whiteboarding and collaboration scenarios.
104
Supported Languages
Chinese, Arabic, Polish, German, Russian, French, Korean, Czech, Romanian, English (US), Afrikaans, Portuguese (PT/BR), Japanese, Thai, Ukrainian, Turkish, Spanish (ES/LATAM/US), Hungarian, Italian, Hebrew, Vietnamese, Indonesian, and 80+ more
$
Delivery & Pricing
Output format: APK (SDK format subject to project-based evaluation). Pricing: NRE (six-to-seven figures RMB range depending on scope) + license at RMB 5 per unit list price.
OVERSEAS READY

Overseas Readiness Summary

OCR SDKOVERSEAS READYoffline, Windows / Android / iOS devices
OCR APIPLANNEDPaddleOCR recommended, 111 languages
Face Offline SDKOVERSEAS READYoffline liveness & recognition
Face APIPLANNEDface search & liveness detection
Offline TTS SDKOVERSEAS READYon-device synthesis, no network
Speech SDK / APIOVERSEAS READYEU & NA nodes proven, per-project evaluation
Translation APIsOVERSEAS READY200+ languages, LLM & classic, custom terms
Input Method (Offline)OVERSEAS READY104 languages, watch / car / TV / large screen

Technical Specifications

Total APIs260+ across 6 categories
Speech Languages20+ including EN, ZH, JA, KO, FR, DE, ES
Translation Pairs100+ language combinations
SDKsPython, Java, Go, Node.js, PHP, C#, C++
API ProtocolREST (JSON) + gRPC + WebSocket (streaming)
Rate LimitsUp to 5,000 QPM (enterprise tier)
Latency (p99)Under 200ms for most vision APIs
SLA99.95% uptime guarantee

Ready to Integrate AI?

Get API keys and start building in minutes

Register Now View Documentation