Georgian is a low-resource language. The digital material that exists is mostly scraped web text and read-aloud sentences — not the spontaneous, conversational speech that speech and language models actually need. We record it, transcribe it, and license it with the paperwork that makes it usable.
Three service lines. All delivered with full metadata, documentation and a signed rights chain.
Monologue and two-speaker dialogue, recorded in a treated room with one microphone and one channel per speaker. Not read speech — the natural, disfluent, overlapping kind that existing Georgian corpora do not contain.
Native-speaker transcription to a documented standard. Fillers and repetitions preserved, non-speech events tagged, personally identifiable information flagged and redacted. Inter-annotator agreement measured and reported.
Preference ranking, evaluation sets and red-teaming in Georgian. Translated preference data does not transfer — idiom, register and cultural reference do not survive machine translation. This work requires native speakers, and we have a vetted network of them.
Default delivery format. We adapt to your pipeline on request — segmentation, sample rate, container and metadata schema are all configurable.
| Parameter | Specification |
|---|---|
| Language | Georgian — ISO 639-3 kat, script Geor (Mkhedruli) |
| Audio format | WAV, 48 kHz, 24-bit, mono — one channel per speaker |
| Processing | None. No noise reduction, gating, compression or normalisation. Raw as recorded. |
| Signal quality | Mean SNR above 30 dB, zero clipping, room tone recorded each session |
| Dialogue | Separate sample-aligned files per speaker; channel separation measured and reported; bleed is retained, never processed out |
| Segment length | Monologue 5–20 s (max 30 s), cut only on silence, never mid-word. Dialogue delivered as whole sessions. |
| Transcription | Verbatim, UTF-8, one transcript per audio file; timestamps optional |
| Speaker diversity | Balanced by gender, age band and region; distribution reported in the data card |
| Documentation | Data card following Datasheets for Datasets, QA report, transcription guidelines, SHA-256 checksums |
| Delivery | Structured archive over secure transfer, or to your specification |
The licence is the product. A clean, well-documented dataset is worth more than a large one with an ambiguous rights chain — and we build for the first case.
What already exists for Georgian, and what does not.
Common Voice (read speech) · FLEURS · FLORES-Plus · Common Crawl derived text · OpenSubtitles · WikiMatrix · localisation strings
Useful as a baseline. Commoditised, and none of it is conversational.
Spontaneous monologue · two-speaker conversation with natural overlap · domain speech (medical, legal, financial) · telephone and far-field conditions · regional accents and dialects · native-speaker preference and evaluation data
None of this can be scraped. It has to be recorded.
We keep a sample package ready — audio, transcripts, full data card and QA report. Tell us what you need and we will send it, or run a paid pilot to your specification.
Or email hello@aiiacorpus.com directly.