Developing an AI-Powered Speech Recognition Engine for a Legal-Tech Company
Success Story

About the Client

client logo

The client is a Nigerian legal technology company focused on streamlining the delivery of judgments. They also provide digital tools that support thousands of law students, legal practitioners, lecturers, and judges in research and case preparation.

Country:

Nigeria

Industry:

Legal

Download PDF

Business Situation and Requirements

Manual transcription in courtrooms required sustained attention to long recordings and precise capture of spoken words. Even small lapses led to missed phrases or incorrect entries in the final transcript. The work was time-intensive and difficult to scale. Variations in accents across Nigeria’s regions made it harder, as transcribers were not always familiar with local speech patterns. This often resulted in incomplete or inconsistent records.

The client wanted a more reliable way to handle transcription. They aimed to build software that could convert courtroom audio into accurate text with minimal manual effort. The focus was on reducing errors, handling diverse accents, and speeding up the process so legal teams could access clear and complete transcripts without delays.

Some of their key requirement included:

  • Develop an AI-powered transcription platform of both streaming audio and video recordings.

  • Train the AI speech model to recognize variations in accents and intonations across Nigeria.

  • Automate data augmentation into formats and sizes consumable by the software.

  • Minimize errors caused by background noise or audio distortions.

  • Continuously maintain and retrain the software using new datasets from court proceedings and judgments.

  • Ensure accurate and error-free transcription to support judges and legal practitioners.

Solution

Our team started by developing an Automated Speech Recognition (ASR) engine using a transfer learning-based Machine Learning (ML) model. This approach allowed the system to learn from existing speech data and adapt quickly to the variety of accents and speaking styles found in Nigerian courts.

The ASR engine was trained to listen to audio in real time and produce accurate transcripts efficiently. The system also includes continuous learning to improve accuracy over time. By automating transcription, the solution reduces manual effort, minimizes errors, and ensures court documents are consistent, reliable, and delivered faster, helping legal proceedings run smoothly.

Some of the key features are as follows:

AI-Driven Data Preparation for Speech Training

  • The ASR model was trained using thousands of hours of audio recordings along with their corresponding transcripts.
  • All audio files were standardized to meet specific requirements, including a mono-channel format and a 16 kHz sampling rate.
  • The data preparation process was automated by the team before feeding it into the ASR engine.
  • This structured approach improved the system’s ability to understand Nigerian accents and local vocabulary.
AI-Driven Data Preparation for Speech Training

Training the Model on 100,000+ Words and Linguistic Samples

  • Prepared the dataset and fed it into the ASR engine for training.
  • Used optional modules for data augmentation and pre-training.
  • Fine-tuned the model using transfer learning to improve accuracy.
  • Included WAV audio files from multiple speakers with different accents.
  • Processed audio files and continuously monitored the model to reduce training loss.
Training the Model

Improved Transcript Readability

  • The ASR model’s post-processing pipeline sometimes produced text with readability issues.
  • To address this, our team implemented an Inverse Text Normalization (ITN) mechanism.
  • This helped refine the output text and improve the clarity of transcribed depositions.

High-Accuracy Transcription

  • The ASR engine can transcribe depositions with near-perfect accuracy.
  • Extensive training helped keep the Word Error Rate (WER) and Character Error Rate (CER) below 20%.
Improved Transcript Readability

Impact

The final solution enabled the legal technology company to help hundreds of courts across Nigeria automate the transcription of case proceedings, hearings, and depositions.

The AI-powered system generated 80% accurate text records of audio depositions, improving the reliability of court transcripts.

This improvement made legal documentation more accurate and helped judges deliver judgments faster. The solution also reduced the documentation turnaround time by 35%. The client highly appreciated Unthinkable’s AI-driven approach and fast execution.

AI-Powered Speech Recognition Engine for a Legal-Tech Company

100,000+

size of training dataset

80%

accurate text records of audio depositions

35%

reduction in documentation processing time