Georgian · ISO 639-3 kat

Georgian language data, collected properly.

Georgian is a low-resource language. The digital material that exists is mostly scraped web text and read-aloud sentences — not the spontaneous, conversational speech that speech and language models actually need. We record it, transcribe it, and license it with the paperwork that makes it usable.

3.7M
Georgian speakers worldwide
kat
Unique script, polypersonal verb morphology
100%
Signed consent, commercial use and sublicensing
48 kHz
24-bit WAV, mono per speaker, unprocessed

What we supply

Three service lines. All delivered with full metadata, documentation and a signed rights chain.

Speech

Spontaneous and conversational audio

Monologue and two-speaker dialogue, recorded in a treated room with one microphone and one channel per speaker. Not read speech — the natural, disfluent, overlapping kind that existing Georgian corpora do not contain.

Transcription

Verbatim transcription

Native-speaker transcription to a documented standard. Fillers and repetitions preserved, non-speech events tagged, personally identifiable information flagged and redacted. Inter-annotator agreement measured and reported.

Annotation

Native-speaker judgment

Preference ranking, evaluation sets and red-teaming in Georgian. Translated preference data does not transfer — idiom, register and cultural reference do not survive machine translation. This work requires native speakers, and we have a vetted network of them.

Technical specifications

Default delivery format. We adapt to your pipeline on request — segmentation, sample rate, container and metadata schema are all configurable.

ParameterSpecification
LanguageGeorgian — ISO 639-3 kat, script Geor (Mkhedruli)
Audio formatWAV, 48 kHz, 24-bit, mono — one channel per speaker
ProcessingNone. No noise reduction, gating, compression or normalisation. Raw as recorded.
Signal qualityMean SNR above 30 dB, zero clipping, room tone recorded each session
DialogueSeparate sample-aligned files per speaker; channel separation measured and reported; bleed is retained, never processed out
Segment lengthMonologue 5–20 s (max 30 s), cut only on silence, never mid-word. Dialogue delivered as whole sessions.
TranscriptionVerbatim, UTF-8, one transcript per audio file; timestamps optional
Speaker diversityBalanced by gender, age band and region; distribution reported in the data card
DocumentationData card following Datasheets for Datasets, QA report, transcription guidelines, SHA-256 checksums
DeliveryStructured archive over secure transfer, or to your specification

Rights and compliance

The licence is the product. A clean, well-documented dataset is worth more than a large one with an ambiguous rights chain — and we build for the first case.

  • Written consent from every speaker, signed before recording, in Georgian with a verified English translation available for audit.
  • Commercial use explicitly granted — not research-only.
  • Sublicensing rights granted, so the data can be passed to your clients and affiliates without a further permission chain.
  • Every participant is paid at rates well above the local market, and all are 18 or older.
  • No scraped material. Everything is recorded for this purpose, first-party, with a documented provenance.
  • PII tagged and redacted in both audio and transcript, with the procedure documented.
  • Consent records retained and verifiable by speaker code on request, without disclosing personal data.
Why this matters commercially. Buyers of training data increasingly need to demonstrate provenance to their own customers, auditors and regulators. A dataset without a documented consent chain is a liability on your balance sheet, not an asset. Everything we deliver is built so that it survives that question.
Voice is biometric data. Under Georgian law voice characteristics fall within the definition of biometric data. We handle collection, storage and notification accordingly, and can provide our data protection documentation during vendor onboarding.

The gap we fill

What already exists for Georgian, and what does not.

Already available publicly

Common Voice (read speech) · FLEURS · FLORES-Plus · Common Crawl derived text · OpenSubtitles · WikiMatrix · localisation strings

Useful as a baseline. Commoditised, and none of it is conversational.

What we produce

Spontaneous monologue · two-speaker conversation with natural overlap · domain speech (medical, legal, financial) · telephone and far-field conditions · regional accents and dialects · native-speaker preference and evaluation data

None of this can be scraped. It has to be recorded.

Request a sample

We keep a sample package ready — audio, transcripts, full data card and QA report. Tell us what you need and we will send it, or run a paid pilot to your specification.

If you have a written spec, mention it here and we will reply with a note on fit and capacity.

Or email hello@aiiacorpus.com directly.