Georgian · ISO 639-3 kat

Georgian language data, collected properly.

Georgian is a low-resource language. The digital material that exists is mostly scraped web text and read-aloud sentences — not the spontaneous, conversational speech that speech and language models actually need. We record it, transcribe it, and license it with the paperwork that makes it usable.

3.7M
Georgian speakers worldwide
kat
Unique script, polypersonal verb morphology
100%
Signed consent, commercial use and sublicensing
48 kHz
24-bit WAV, mono per speaker, unprocessed

Hear a sample

Thirty seconds of spontaneous Georgian, unprocessed, with the verbatim transcript. This is the raw signal as delivered — no noise reduction, no gating, no normalisation.

0:00 / 0:00

What we supply

Three service lines. All delivered with full metadata, documentation and a signed rights chain.

Speech

Spontaneous and conversational audio

Monologue and two-speaker dialogue, recorded in a treated room with one microphone and one channel per speaker. Not read speech — the natural, disfluent, overlapping kind that existing Georgian corpora do not contain.

Transcription

Verbatim transcription

Native-speaker transcription to a documented standard. Fillers and repetitions preserved, non-speech events tagged, personally identifiable information flagged and redacted. Inter-annotator agreement measured and reported.

Annotation

Native-speaker judgment

Preference ranking, evaluation sets and red-teaming in Georgian. Translated preference data does not transfer — idiom, register and cultural reference do not survive machine translation. This work requires native speakers, and we have a vetted network of them.

Technical specifications

Default delivery format. We adapt to your pipeline on request — segmentation, sample rate, container and metadata schema are all configurable.

ParameterSpecification
LanguageGeorgian — ISO 639-3 kat, script Geor (Mkhedruli)
Audio formatWAV, 48 kHz, 24-bit, mono — one channel per speaker
ProcessingNone. No noise reduction, gating, compression or normalisation. Raw as recorded.
Signal qualityMean SNR above 30 dB, zero clipping, room tone recorded each session
DialogueSeparate sample-aligned files per speaker; channel separation measured and reported; bleed is retained, never processed out
Segment lengthMonologue 5–20 s (max 30 s), cut only on silence, never mid-word. Dialogue delivered as whole sessions.
TranscriptionVerbatim, UTF-8, one transcript per audio file; timestamps optional
Speaker diversityBalanced by gender, age band and region; distribution reported in the data card
DocumentationData card following Datasheets for Datasets, QA report, transcription guidelines, SHA-256 checksums
DeliveryStructured archive over secure transfer, or to your specification

Capacity and turnaround

What we can commit to today. Larger volumes are possible with lead time — tell us the number and we will tell you honestly whether we can meet it.

40–60 h
Audio hours per month, scalable with notice
10 days
Typical pilot turnaround from signed brief
5 hours
Minimum engagement
Your spec
Format, segmentation and metadata schema

How we work

Four steps. No long procurement cycle before you can judge the quality.

Step 1

Brief

You send the specification — volume, speech type, speaker distribution, format. We reply with a note on fit and whether we can meet it.

2 business days
Step 2

Pilot

A small batch built exactly to your spec, with the full documentation set. You evaluate real data before committing to volume.

10 business days
Step 3

Production

Delivered in batches, not one drop at the end. Each batch carries its own QA report so problems surface early, not on the final handover.

Rolling
Step 4

Handover

Structured archive with checksums, data card and consent documentation. Signed consent records remain with us and stay verifiable by speaker code.

On completion

What a delivery looks like

Every package has the same shape. Nothing is assembled by hand at the end, so nothing is forgotten.

georgian-speech-v1/ │ ├── README.md what this is, how to use it ├── DATA_CARD.md full datasheet — composition, collection, limitations ├── LICENSE.txt exact rights granted ├── metadata.csv one row per recording, 18 fields ├── checksums.txt SHA-256 for every file │ ├── audio/ WAV · 48 kHz · 24-bit · mono per speaker ├── transcripts/ UTF-8, one per audio file │ └── docs/ ├── transcription_guidelines.md ├── qa_report.md SNR distribution, IAA, rejected files └── consent_template.pdf blank form, Georgian and English
Why the QA report ships with the data. It states the measured signal quality, the inter-annotator agreement and which files were rejected and why. You should not have to take our word for the quality — the numbers come in the box.

Rights and compliance

The licence is the product. A clean, well-documented dataset is worth more than a large one with an ambiguous rights chain — and we build for the first case.

  • Written consent from every speaker, signed before recording, in Georgian with a verified English translation available for audit.
  • Commercial use explicitly granted — not research-only.
  • Sublicensing rights granted, so the data can be passed to your clients and affiliates without a further permission chain.
  • Every participant is paid at rates well above the local market, and all are 18 or older.
  • No scraped material. Everything is recorded for this purpose, first-party, with a documented provenance.
  • PII tagged and redacted in both audio and transcript, with the procedure documented.
  • Consent records retained and verifiable by speaker code on request, without disclosing personal data.
Why this matters commercially. Buyers of training data increasingly need to demonstrate provenance to their own customers, auditors and regulators. A dataset without a documented consent chain is a liability on your balance sheet, not an asset. Everything we deliver is built so that it survives that question.
Voice is biometric data. Under Georgian law voice characteristics fall within the definition of biometric data. We handle collection, storage and notification accordingly, and can provide our data protection documentation during vendor onboarding.

The gap we fill

What already exists for Georgian, and what does not.

Already available publicly

Common Voice (read speech) · FLEURS · FLORES-Plus · Common Crawl derived text · OpenSubtitles · WikiMatrix · localisation strings

Useful as a baseline. Commoditised, and none of it is conversational.

What we produce

Spontaneous monologue · two-speaker conversation with natural overlap · domain speech (medical, legal, financial) · telephone and far-field conditions · regional accents and dialects · native-speaker preference and evaluation data

None of this can be scraped. It has to be recorded.

Request a sample

We keep a sample package ready — audio, transcripts, full data card and QA report. Tell us what you need and we will send it, or run a paid pilot to your specification.

If you have a written spec, mention it here and we will reply with a note on fit and capacity.

Or email hello@aiiacorpus.com directly.