Multilingual NLP, Model Evaluation & Document AI

Hicham YASSIN | Multilingual NLP & Model Evaluation Engineer

  • AI training and model evaluation contributor at Scale AI • Designed and reviewed prompts and model responses for evaluation-oriented workflows.
  • M.Sc. in Computer Science (NLP & Smart Computing) • Nantes Université, with research work in historical HMER.
  • French, English and Arabic • Arabic-French translation, terminology work, linguistic data curation and editorial quality control.
  • LLM/VLM adaptation and output analysis • Python ML engineering, prompt design and reproducible evaluation on historical document tasks.
PRESERVATION & TRANSMISSION Classical Knowledge & Didactics

"Languages are the best mirror of the human mind, and a precise analysis of the meaning of words would disclose the operations of the understanding better than anything else."

— Leibniz, New Essays on Human Understanding

Arabic-French Translation

Specialized translation, terminology choices and bilingual editorial quality control for scholarly texts.

LLM/VLM Adaptation

Fine-tuning and evaluation workflows for historical handwritten manuscripts and mathematical notation.

Model Output Analysis

Prompt design, response review and error analysis across reasoning, structure and transcription quality.

Python ML Engineering

Data processing, evaluation scripts, reproducible experiments and practical web interfaces.

Research & Development: Multilingual NLP, Document AI & Evaluation

Applied modeling for historical manuscripts, multilingual linguistic data, model output analysis and practical AI systems.

Python R&D

HTR for Historical Maghrebi Manuscripts (Qwen 9B)

Supervised fine-tuning (LoRA/QLoRA via **Unsloth**) of Qwen 3.5 9B on a historical cursive manuscript dataset. Target CER: 13%. (Private model, planned for open-source release).

Maghrebi Dataset Distribution
Total Volume ~11K annotated lines
Data Splits
  • Train: 9,524 lines (85%)
  • Val: 560 lines (5%)
  • Test: 1,121 lines (10%)
Source Distribution
  • RASAM I & II: 8,444 lines
  • BULAC (Suppl.): 962 lines
  • Custom Annotations: 931 lines
  • BnF (Suppl.): 581 lines
  • MMSH Archives: 287 lines
Modeling & R&D Axes
  • Masked Autoencoders (MAE): ViT encoder initialization via masked pixel reconstruction (He et al., 2021) to stabilize vision modeling against paper degradation.
  • Crowdsourcing: Continuous ingestion of verified annotations from Khatt-GPT.
  • Segmentation & HTR: Interactive online demo available at htr.omnignosis.fr/manuscrit.
Transcription Visualizer (Maghrebi HTR)
Original Maghrebi Manuscript
و الخامس ما اضيف الى واحد من هذه الاربعة المذكورة تقول في Mضاف الى المضمر غلامي
Transcribed (Qwen-VL LoRA - 13% CER)
Slide to simulate automatic classical Arabic transcription
PHILIUMM - Nantes Université / LS2N

Structure-Aware Recognition of Historical Mathematical Expressions

Historical HMER on Leibniz manuscripts using CoMER, structural supervision and QLoRA-adapted vision-language models.

Research on historical handwritten mathematical expression recognition using the Leibniz-HME corpus. The study compares a compact CoMER recognizer with a 4-bit QLoRA adaptation of Uni-MuMER-Qwen3.5-2B under the fixed Spatial-Constraint evaluation protocol. Manuscript in progress, not yet published.

73.05% ExpRate - adapted Qwen
0.95 Mean token edit distance
961 Held-out test expressions
43.27% Hard-expression accuracy
Full context Leibniz scan Full context image (original manuscript)
Cropped Leibniz expression Corresponding cropped mathematical expression
Leibniz HMER Adaptation & Evaluation Protocol
  1. Corpus: Leibniz-HME contains 4,571 annotated mathematical expressions from 85 manuscript pages with a normalized 140-token vocabulary.
  2. Protocol: Results use the fixed Spatial-Constraint split: 2,642 training, 968 validation and 961 test expressions, using pre-localized expression crops.
  3. CoMER: Adapted a compact recognizer with external HMER pretraining and auxiliary Symbol Counting and Tree-CoT supervision.
  4. Qwen adaptation: Adapted the public Uni-MuMER-Qwen3.5-2B checkpoint with 4-bit QLoRA and PLAIN, Symbol Counting, Tree-CoT and Error-Driven Learning supervision.
Quantitative Results (Leibniz S-C Test Split)
Model & Alignment Recipe ExpRate (%) Edit Distance
CoMER + external HMER pretraining + Tree-CoT 72.01 % 1.12
CoMER + Symbol Counting + Tree-CoT 72.01 % 1.14
Uni-MuMER-Qwen3.5-2B + 4-bit QLoRA + PLAIN/SC/Tree-CoT/EDL 73.05 % 0.95
ExpRate is exact expression recognition rate on the fixed S-C test split. Reference-aware oracle analysis is not presented here as deployable model performance.
Scientific Contributions & Conclusions
  • Historical adaptation: Compared compact CoMER recognition with QLoRA adaptation of Uni-MuMER-Qwen3.5-2B on historical notation.
  • Structural supervision: Used Symbol Counting and Tree-CoT objectives to expose omissions, grouping errors, decorators and rare historical symbols.
  • Error analysis: Analyzed residual failures across structural scope, missing symbols and notation-specific ambiguities.
  • EDL caution: Conditional EDL repaired 27 of 259 rejected PLAIN predictions exactly, but was not reliable as an unconditional post-processing system.
Alignment Visualizer (Leibniz HMER)
Original Leibniz Manuscript
y \gleich \frac{x}{x + 1} Example LaTeX overlay - qualitative visualization
Slide to align the original manuscript formula with its LaTeX transcription
Work in Progress | Arabic OCR

Arabic-OCR: OCR-to-Translation Dataset Pipeline

Current project focused on scanned Arabic manuscripts: preprocessing, OCR experiments, translation alignment and dataset tooling.

OCR Arabic manuscript scans
Translation Arabic-French alignment
Datasets Curation workflows
  • Building image preprocessing and OCR evaluation workflows for Arabic manuscript material.
  • Connecting OCR outputs to translation review and bilingual dataset creation.
  • Project in active development; no public demo link yet.
Next.js 16+ | Crowdsourcing

Khatt-GPT: Collaborative HTR Platform

Web crowdsourcing platform for collaborative transcription of historical Maghrebi manuscripts. Streamlines pedagogical annotation ingestion to enrich custom datasets.

khattgpt.vercel.app Platform Link
Crowdsourcing Assisted Annotation
Next.js 16+ React Dataset Collector TailwindCSS Vercel Serving
Hackathon Riyadh | SDAIA

AjurrumAI: Grammar Pedagogy

Interactive application for teaching classical Arabic grammar (I'rab) and translation (SDAIA Saudi Hackathon, Riyadh). Fine-tuned Saudi ALLAM LLM (IBM Watsonx.ai) on proprietary datasets.

ALLAM LLM IBM Watsonx.ai Arabic NLP I'rab Dataset SDAIA Saudi
Hackathon Llama 3 | lablab.ai

Arabic Companion: Adaptive Language Learning

Personalized conversational agent for learning classical Arabic. Features real-time grammar feedback, correction loops, and adaptive vocabulary learning. Developed at the Llama 3 Hackathon.

Llama 3 LlamaIndex Together AI LangChain Arabic NLP
Next.js 16+ | Workspace

HTR & Speech Transcription Workspace

Collaborative HTR workspace (line segmentation, visual alignment aids) and speech-to-text transcription (Whisper STT).

Next.js 16+ Line Segmentation Speech-to-Text Whisper / Translation Local LLM API
Didactics & Humanities

Arabic-French Translation & Heritage Publishing

Specialized Arabic-French translation, terminology control, bilingual linguistic resources and editorial quality workflows.

  • Founder of Héritage Mohammadien, an independent publishing house for Arabic-French books, translation and editorial production.
  • Translate, write and publish books; handle typesetting, cover design, terminology choices and source-based editorial quality control.
  • Developed bilingual digital publishing, OCR-to-translation and data-processing workflows for fr.mahdara.org and related materials.

Bridging classical humanities, translation quality and rigorous language-data workflows.

Arabic-French Translation Terminology QA Bilingual Portals Linguistic Data Curation
Select classical books translated, annotated, and published by Hicham YASSIN.

Hicham YASSIN

Multilingual NLP & Model Evaluation Engineer

French, English and Arabic - model evaluation, linguistic data, LLM/VLM adaptation and document AI

Email: Mathis.yassin@gmail.com

Phone: +33 7 63 26 26 67

Location: Nantes, France & Rabat, Morocco

GitHub: github.com/h-yassin-ai

Websites: heritagemohammadien.fr | fr.mahdara.org | khattgpt.vercel.app

Multilingual NLP engineer working across French, English and Arabic. Builds evaluation workflows, OCR/HTR pipelines, datasets and web tools, with a research background in NLP and document AI.

Education

Université de Nantes, France
M.Sc. in Computer Science — NLP & Smart Computing (ATAL)

NLP, deep learning, Transformers, speech processing, computer vision and reinforcement learning.

Université du Mans, France
Double B.Sc. Major in Mathematics & Economics

Mathematics, statistics, optimization, econometrics and quantitative modeling.

Professional Experience in AI

May 2025 - Aug 2025 Scale AI - Remote
AI Training and Model Evaluation Contributor
  • Reviewed prompts and model outputs for AI training/evaluation.
  • Checked reasoning, instruction-following, response structure and output quality.
  • Applied QA rubrics in iterative review workflows.
Independent Projects
Freelance Web Developer & RAG Chatbot Integrator
  • Built websites, internal tools and AI assistants with React/Next.js, Python/FastAPI and API integrations.
  • Integrated RAG chatbots for document search, retrieval workflows and domain-specific Q&A.
  • Delivered end-to-end projects from client need to deployment, content structure and maintenance.

Research Projects & AI Platforms

Khatt-GPT
Collaborative HTR Platform (khattgpt.vercel.app)
  • Built a web platform for manuscript transcription, annotation and HTR dataset curation.
PHILIUMM Project - Leibniz HMER
Structure-Aware Historical HMER - Leibniz Manuscripts
  • Adapted CoMER and Uni-MuMER-Qwen3.5-2B with QLoRA and structural supervision.
  • Reached 73.05% ExpRate and 0.95 token edit distance on the 961-expression S-C test split.
  • Research codeDemo code
HTR Modeling (Private Project)
Historical Maghrebi HTR - Qwen 3.5 9B
AjurrumAI (ALLAM & SDAIA Hackathon, Riyadh)
AjurrumAI: Arabic Grammar Pedagogy
  • Fine-tuned ALLAM for Arabic syntax and translation; built a grammar UI with syntax trees.
July 2024 Arabic Companion (Llama 3 Hackathon, lablab.ai)
Arabic Companion: Conversational Language Tutor
  • Built an Arabic learning assistant with Llama 3, LlamaIndex, Together AI and LangChain.
Arabic-OCR
Arabic OCR Pipeline - Work in Progress
  • Building scanned-manuscript preprocessing, OCR-to-translation experiments and dataset tooling.

Technical Expertise

I. NLP, Evaluation & Model Adaptation
Multilingual NLP Model evaluation Prompt design Linguistic data curation Transformers & VLMs Fine-tuning (LoRA/QLoRA) PyTorch & Hugging Face HTR & Computer Vision
II. Software & Web Engineering
Next.js 15/16 TypeScript (Turborepo) Python FastAPI Docker RAG chatbots Git
III. Infrastructure & Tooling
Linux Administration VPS (Coolify) Local Inference Data pipelines
IV. Languages & Publishing
French (Native) English (C1) Literary Arabic (Advanced) Arabic-French translation

Translation & Dataset Work

Éditions Héritage Mohammadien (Since 2019)

Translate, write and publish Arabic-French books through an independent publishing house. Handles typesetting, cover design, web publishing and AI-assisted OCR-to-translation dataset work.

Languages

  • French: Native
  • English: Professional (C1)
  • Classical Arabic: Scholarly literacy and oral comprehension

Interests

Chess & Game Theory Logic Puzzles Typography & Medieval Bookbinding Philosophy of Science Pedagogical Transmission