DocumentationIntegrations

Rust · 0.1.0

Audio transcription

Configure a supported provider and understand what is sent outside the instance.

Configure a provider

Administrators can configure OpenAI or Gemini provider settings and select a transcription provider/model. The Rust API exposes Transcribe for authenticated users. The implementation accepts audio content bytes, not an audio URI, and rejects empty audio or content larger than 25 MiB. Provider-specific limits can be stricter.

Consent and scope

Transcription sends the submitted audio to the configured external provider. Confirm you are allowed to share that recording and understand the provider's data policy and possible charges. Store provider keys in authorized settings or deployment secrets, never in public examples. This documented feature is audio transcription; it does not promise automatic memo tagging, summarization or chat.

Configure provider and transcription separately

As an administrator, open Settings → AI and add a provider. Supply a descriptive title, choose OPENAI or GEMINI, and enter the provider key. A blank endpoint selects the built-in provider endpoint; use a custom endpoint only if you trust it to receive the key and audio. Save the provider. In the transcription section, select that provider, choose a model, and save transcription separately. Language and prompt are optional hints. Saving a provider does not automatically select it for transcription.

Submit a small disposable recording

This example needs Python 3, curl, a short WAV file named sample.wav, and a configured authenticated instance. Protobuf JSON encodes bytes as base64. The request contains audio.content, filename and contentType; URI inputs are not implemented. The Python step creates only a local request file; the curl step sends the recording to your instance and then to its configured provider, potentially incurring charges. Use only a recording you are allowed to transmit.

umask 077
python3 - <<'PYTHON'
import base64, json
from pathlib import Path
request = {"audio": {
    "content": base64.b64encode(Path("sample.wav").read_bytes()).decode(),
    "filename": "sample.wav",
    "contentType": "audio/wav"
}}
Path("transcription-request.json").write_text(json.dumps(request))
PYTHON
curl --fail-with-body "$MEMOS_URL/api/v1/ai:transcribe" \
  -H "Authorization: Bearer $MEMOS_TOKEN" \
  -H 'Content-Type: application/json' \
  --data-binary @transcription-request.json

Check the transcript and known limits

A successful response has a text field containing the transcript; it does not automatically create a memo. Read and correct the text before saving it. Delete the temporary request file when no longer needed because it contains the audio in reversible base64. “Transcription is not configured” means no saved transcription provider is selected. The outer request permits at most 25 MiB of decoded audio; Gemini additionally limits inline content to 14 MiB after any WebM-to-WAV conversion and accepts a narrower set of formats. For Gemini, start with a small audio/wav or audio/mp3 file. An unsupported MIME type, provider HTTP error or incomplete provider response requires correcting that input/configuration, not repeatedly submitting the same recording.