Skip to content
View SethBastianUade's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SethBastianUade

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SethBastianUade/README.md
Sebastián Arroyo — Building Reliable Software. Studying Reliable AI.

Portfolio  ·   GitHub  ·   LinkedIn  ·   Email

Mission

This repository is a public log of my transition from backend engineering into AI safety and AI engineering.

I spend my days building reliable software — Java, Spring Boot, well-shaped APIs — and my evenings studying what it would take to build reliable AI. I document the work in the open: notes, reading, small experiments, and the things I get wrong along the way.

The intended output is not a portfolio. It is a long-running notebook, kept for years, that records how an engineer actually moves into this field — one paper, one eval, one failed reproduction at a time.


Research Timeline

A working history. Intentionally incomplete — milestones get added as they happen.

2024 ──┐
       │  Started Software Engineering
       │
2025 ──┤
       │  Professional Java Backend Engineer
       │
2026 ──┤
       │  Selected for BlueDot Technical AI Safety
       │
Now  ──┤
       │  Building Kreta
       │
Now  ──┤
       │  Learning AI Safety · Alignment · LLM Reliability
       │
Next ──┤
       │  AI Alignment Engineering
       │
...   └─ future entries

Research Log

A live laboratory notebook. Each topic is a working file under /research. Most entries are preliminary — they grow as understanding does.

  Status legend:  Reading active study · Drafting summary stable · Working running experiments · Revisited returned to after a gap

Status Topic Summary Notes
Reading Prompt Injection Instruction/data boundary inside LLM systems placeholders
Reading Mechanistic Interpretability Reverse-engineering internal model computation placeholders
Reading AI Evaluations Capability and safety evals as infrastructure placeholders
Reading Transformer Circuits Attention and MLP circuits, induction heads placeholders
Reading Agent Safety Failure modes of autonomous LLM loops placeholders
Reading Scalable Oversight Supervising stronger actors from weaker ones placeholders
Reading Constitutional AI Self-critique against written principles placeholders

Topics are added as I start them. Files I have not yet opened are not listed here.


Current Reading

A queue, tracked honestly. Bars show where I am; links go to notes when they exist.

AI Safety Fundamentals  — BlueDot curriculum ██████████░░░░░░░░░░   50%  ·  Reading  ·  notes pending

Mathematical Framework for Transformer Circuits  — Elhage et al. ████░░░░░░░░░░░░░░░░   20%  ·  Reading  ·  notes pending

In-context Learning and Induction Heads  — Olsson et al. ██░░░░░░░░░░░░░░░░░░   10%  ·  Queued

Language Models Can Explain Neurons in Language Models  — OpenAI ░░░░░░░░░░░░░░░░░░░░   0%  ·  Queued

Constitutional AI  — Bai et al., Anthropic ██░░░░░░░░░░░░░░░░░░   10%  ·  Queued

backlog — papers queued for later (click to expand)
  •  Toy Models of Superposition — Elhage et al.
  •  Sleeper Agents — Hubinger et al.
  •  Scaling Monosemanticity — Templeton et al.
  •  AI Safety via Debate — Irving et al.
  •  Are Emergent Abilities a Mirage? — Schaeffer et al.
  •  Not what you've signed up for — Greshake et al.
  •  placeholder — add next paper

Experiments

Small implementations to test understanding. Most fail or produce noise. That is the point.

Status Name What it tests Repo
Planning llm-eval-001 A minimal capability + safety eval harness from scratch coming
Planning induction-heads Reproduce induction-head detection on GPT-2 small coming
Planning inject-lab A controlled indirect-prompt-injection testbed against a tool-using agent coming
Planning weak-to-strong Minimal weak-to-strong generalisation experiment coming
Planning agent-harness An LLM agent with an explicit, auditable permission surface coming

Empty repos are not created until I have time to start them. Names are reserved.


Lab Notes

Short, dated entries. Easy to maintain by hand. Newest on top.

  2026-07-17  — Started this notebook publicly.   Kreta continues in parallel; research log begins here.

  2026-07-15  — Accepted into BlueDot Technical AI Safety.

  2026-07-  — placeholder


Open Questions

Things I cannot answer yet and am actively thinking about. Listed to show what I am working toward, not what I claim.

  • How should we evaluate reasoning rather than retrieval?
  • What practical interpretability techniques transfer from small to frontier models?
  • How can agent failures — long-horizon, sparse, compounding — be measured before deployment?
  • Is instruction/data separation a property we can train, or only scaffold?
  • What is the smallest meaningful eval set that predicts deployment risk?
  • Can oversight keep up when the actor outpaces the overseer?
  • How do we detect when a model has internalised a principle versus learning to recite it?

Added as I notice them. Never removed; downgraded to a Research Log entry when answered.


Focus Areas

Engineering
Java · Spring Boot · System Design · Enterprise APIs

·  Backend services with clear architecture
·  Maintainable, observable systems
·  Developer tooling and DX
Research
AI Safety · Alignment · LLM Reliability

·  Reliable AI systems
·  Model evaluations
·  Mechanistic interpretability
Building
Kreta — home for modern software and AI products

Early-stage work where backend reliability meets applied AI.
Learning
BlueDot — Technical AI Safety

Structured study of fundamentals, alignment, and frontier risks.

Stack

Java  Spring Boot  Python  Docker  PostgreSQL  GitHub Actions  React  TypeScript  OpenAI API  Claude Code  Inspect AI

A notebook kept openly. Continuous learning, not finished work.

Popular repositories Loading

  1. Proyect-Club-Center-Fight-Academy Proyect-Club-Center-Fight-Academy Public

    Pagina web en HTML Y CSS, con futuro en backend en JAVA y SQL con springboot

    HTML

  2. SethBastianUade SethBastianUade Public

    Config files for my GitHub profile.

  3. final_proyect_UADE_1 final_proyect_UADE_1 Public

    Final proyect in python to Introducción a la Programación from UADE

    Python

  4. Kash Kash Public

    Proyecto para la materia Algoritmos y estructura de datos 1. Objetivo: Realizar una billetera virtual, donde el usuario pueda vincular multiples cuentas bancarias, realizar transferencias y generar…

    Python 2

  5. UADE-2025-PROGRA-II UADE-2025-PROGRA-II Public

    Proyecto de Programación II 2025

    Java

  6. SethBastianUade.github.io SethBastianUade.github.io Public

    CV Sebastian Arroyo BackEnd Java

    TypeScript