<AI_Engineer_&_ML_Specialist>

/* AI Engineer specializing in RAG architectures, document intelligence, and LLM fine-tuning. Experienced in building enterprise-grade data extraction pipelines and production-ready APIs */

🟢 Open to work
username.png
Currently an Extern @Pfizer

<Work_samples>

Projects include AI models and pipelines for extracting structured data and insights from documents, prototyped during coursework and a Pfizer externship.

  • Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship
    [01]Pfizer Externship⏱️ In progress
    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    ✅ Verified by Extern

    Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

    Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

    AI & MLPythonDocument IntelligencePresentation Skills
  • NBA Discourse Classifier
    [02]Freelance Project
    NBA Discourse Classifier

    Fine-tuned a DistilBERT sequence classification model to evaluate community discourse quality by categorizing social media posts into Analysis, Hot Take, or Reaction classes. Developed a comparative benchmarking pipeline evaluating the fine-tuned models performance and inference latency against a zero-shot Large Language Model (LLM) baseline. Preprocessed and tokenized text datasets using Hugging Face pipelines, optimizing training hyperparameters in PyTorch to maximize classification accuracy.

    PythonPyTorchHugging Face (DistilBERT)TransformersJupyter

<About_me>

AI Engineer specializing in RAG architectures, document intelligence, and LLM fine-tuning. Experienced in building enterprise-grade data extraction pipelines and production-ready APIs

I am Oscar Cabezas, a senior Computer Science student at the University of New Orleans focused on artificial intelligence. I have some practical experience and I am working to build stronger technical skills, especially in AI for document analysis and data extraction.

Externships

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

Pfizer

Experience

Senior AI Data Specialist

iMerit Technology · Jan. 2018 10Present

Freelance Geospatial Annotation Specialist

Pathr.ai · Apr. 2023 02Dec. 2025

Undergraduate Machine Learning Research Assistant

University of New Orleans · Mar. 2024 04Aug. 2024

Education

LSU New Orleans (formerly University of New Orleans)

Bachelor of Science in Computer Science, Concentration in Artificial Intelligence and Machine Learning · Class of 2026

CodePath AI Engineering

Applications of AI Engineering & Foundations of AI Engineering · Class of 2026

Delgado Community College

Associates in Business Administration, Graduated with Honors · Class of 2022

Skills

Artificial intelligenceDocument data extractionMachine learning modelsPythonData preprocessingModel evaluation

✅ Verified by Extern · ⏱️ In progress

<Pfizer_Advanced:_AI-Powered_Document_Insights_&_Data_Extraction_Externship>

Prototype AI-powered document intelligence with Pfizer—using OCR, LLMs, and RAG to automate real enterprise PDF workflows and build a standout portfolio project.

AI & MLPythonDocument IntelligencePresentation Skills

/Overview

The project prototyped AI workflows for extracting and structuring data from enterprise PDF documents, combining OCR, large language models, and retrieval-augmented generation. The work surveyed LLM architectures and OCR pipelines, designed a processing pipeline, and produced prototype components for document ingestion and question answering.

Pfizer Advanced: AI-Powered Document Insights & Data Extraction Externship

//What I've accomplished

I produced a written summary explaining how LLMs tokenize input, use Transformer self-attention to model token relationships, are pre-trained on large corpora, and are fine-tuned with instruction tuning and human feedback.

///Project breakdown

The submission summarized how LLMs tokenized text, used Transformer self-attention to model token relationships, were pre-trained on large corpora, and then fine-tuned with instruction tuning and human feedback to improve responses.

View_all_work

<NBA_Discourse_Classifier>

Fine-tuned a DistilBERT sequence classification model to evaluate community discourse quality by categorizing social media posts into Analysis, Hot Take, or Reaction classes. Developed a comparative benchmarking pipeline evaluating the fine-tuned models performance and inference latency against a zero-shot Large Language Model (LLM) baseline. Preprocessed and tokenized text datasets using Hugging Face pipelines, optimizing training hyperparameters in PyTorch to maximize classification accuracy.

PythonPyTorchHugging Face (DistilBERT)TransformersJupyter
View_all_work

<Provenance_Guard_0_Multi-Signal_Content_Attribution_API>

Architected a resilient Flask API using a multi-signal ingestion pipeline to classify text submissions via parallel semantic analysis and native Python stylometric engines (MATTR and Sentence Length CV). Integrated Groq LLM (Llama-3.3-70b-versatile) for semantic pacing evaluations and engineered a Disagreement Override algorithm to shield creators from false positives by routing divergent signals to an Uncertain safety layer. Designed a transaction-safe SQLite persistence layer governed by a strict state-machine workflow (classified, under review, upheld/overturned) and configured Flask-Limiter IP tracing gates (10 req/min).

FlaskGroq LLMPythonSQLite
View_all_work