[ 00 ] — AI Engineering · Data Architecture

Working AI systems.
Not slides.

We take GenAI pilots to production: agentic systems and RAG, LLM observability, MLOps. Behind the delivery stands well over a decade in Data Science / ML — we work reproducibly and auditably, and the IP stays on your side.

[ 01 ] — SERVICES

Five ways we deliver — plus an ongoing engagement.

A named scope, a fixed timeframe, a concrete artifact at the end. We give a fixed quote after a 30-minute scoping call.

PAKIET 01 · 4–6 weeks

RAG / Agent Proof-of-Value

For companies that want proof of value before investing in a platform.

You get
  • A working demonstrator on your data
  • A response-quality evaluation dashboard
  • A go/no-go decision report

We define the business metric right at the start — so your project doesn’t join the vast majority of GenAI pilots that show no measurable return (MIT). If the results don’t defend the metric, you get an honest report: why, and what’s next.

Get in touch →
PAKIET 02 · 2–3 weeks

LLM Observability Sprint

For teams running LLMs in production with no visibility into quality and cost.

You get
  • Observability (Langfuse / OpenTelemetry)
  • LLM-as-a-Judge evaluation with an audit trail
  • Cost per request

An audit trail and evaluation records for EU AI Act requirements. Transparency obligations (Art. 50) have applied since 08.2026; requirements for high-risk systems were moved by the digital omnibus to 12.2027 (Annex III) and 08.2028 (Annex I). The technical groundwork takes longer to build than that runway suggests. We deliver it — we leave interpreting the regulations to lawyers.

Get in touch →
PAKIET 03 · 2 weeks
Most common starting point

AI Architecture Blueprint

For boards and CTOs ahead of an investment decision.

You get
  • A target-architecture document
  • An implementation backlog
  • A TCO estimate

An independent audit: you get the document, the backlog and the estimate — you can build with anyone. We don’t upsell ourselves an implementation “on the side”.

Get in touch →
PAKIET 04 · from 3 weeks

PoC-to-Production

For DS teams with notebooks that need to deliver repeatably.

You get
  • Reproducible pipelines (Dagster / SQLMesh)
  • Data validation and CI/CD for ML
  • Full auditability

More than half of AI pilots never reach production (Gartner). We close that gap: from notebook to a reproducible, auditable pipeline.

Get in touch →
PAKIET 05 · 3–5 weeks

Entity Resolution & Knowledge Graph Sprint

For companies whose supplier, product, customer and technology dictionaries have drifted apart across systems.

You get
  • Master-dictionary deduplication and canonicalization — from deterministic to semantic matching
  • A controlled vocabulary: decision-grade definitions, scope notes, a relation typology
  • A graph schema (property graph or SKOS/RDF) with structural validation in CI

“Cleaned up” is not a claim — it’s a validator’s output: a report of merged and rejected pairs with rationales, plus structural tests in CI that keep guarding quality after we leave.

Get in touch →
PAKIET 06 · retainer, 2–4 days / month

AI Engineer on your team

For companies that need a senior GenAI engineer on an ongoing, flexible basis.

You get
  • Ongoing access to a senior AI engineer
  • Architecture and code reviews
  • Team mentoring

Max. 2–3 clients in parallel — exclusivity instead of agency scale. Piotr builds himself, not just advises.

Get in touch →

Not sure where to start? Book 30 minutes →

[ 02 ] — FOR WHOM

Who we work with

01

Mid-market and scale-ups with a stalled pilot

You have a GenAI PoC that hallucinates or doesn’t deliver ROI. We take it to production — or tell you plainly that it isn’t worth pursuing, and why.

02

Software houses and integrators

You added AI to your offering, but you lack a senior GenAI engineer at production level. We bring the AI layer under your brand — project subcontracting, not staff leasing.

03

Companies with high-risk systems (EU AI Act)

HR-tech, scoring, fintech — you need the technical groundwork for compliance: logging, documentation, evaluation and an audit trail.

[ 03 ] — PROCESS

Incremental. Payment on milestone acceptance.

  1. 01

    Talk

    A free consultation. We define the problem and the success criteria.

  2. 02

    Blueprint

    Architecture, milestones, estimate. You know what you’re paying for.

  3. 03

    Build

    Every stage ends in a working product. You accept it — you pay.

  4. 04

    Handover

    The IP, documentation and reproducible pipelines stay with you.

[ 04 ] — STACK

Technology stack

Agentic systems in production
  • Claude Agent SDK
  • MCP — production servers
  • Agent Skills (SKILL.md)
  • LangGraph 1.0
  • multi-agent systems
  • PydanticAI v2
  • Temporal — durable execution
  • agent sandboxing (Firecracker / E2B)
  • A2A / AG-UI protocols
Context engineering & RAG
  • agentic retrieval
  • RAG: hybrid dense + BM25 + reranking
  • prompt engineering
  • vision retrieval (ColPali / ColQwen)
  • LazyGraphRAG
  • context compaction & isolation
  • Docling + Mistral OCR
  • deep-research loops
  • embeddings & rerankers (Voyage, Qwen3-Embedding)
Knowledge graphs & agent memory
  • Graphiti — bi-temporal memory
  • Neo4j + GQL (ISO/IEC 39075)
  • Cypher / Text2Cypher
  • Splink — entity resolution
  • LinkML
  • SKOS
  • RDF 1.2 / SPARQL 1.2
  • SFIA / ESCO / O*NET
  • NetworkX
Evals, red-teaming & compliance (EU AI Act)
  • Inspect AI (UK AISI)
  • DSPy 3 + GEPA
  • calibrated LLM-as-a-Judge
  • multi-model consensus
  • faithfulness gates (Ragas)
  • model promotion: shadow / canary
  • red-teaming: promptfoo / garak / PyRIT
  • OTel GenAI + Langfuse (self-hosted)
  • Presidio — PII
  • Annex IV / ISO/IEC 42001
Models, cost & sovereignty
  • frontier behind eval gates (Claude / GPT-5 / Gemini)
  • LiteLLM / Bifrost — gateway
  • cost & latency: cache (Redis), p95 budgets
  • Mistral Large 3
  • Qwen3 + SLM-first
  • LoRA / QLoRA — SLM fine-tuning
  • Hugging Face + Ollama — local models
  • Bielik / PLLuM
  • vLLM + XGrammar
  • classical ML where it beats an LLM
  • VPC / on-prem EU deployments
Data architecture — lean lakehouse
  • PostgreSQL (pgvector / ParadeDB)
  • DuckDB + DuckLake
  • dlt — ingestion as code
  • SQLMesh / dbt (Fusion)
  • Dagster
  • Qdrant — self-hosted EU
  • Lance — multimodal storage
  • Polars
  • uv / ruff / pytest + hypothesis
Baseline — listed for completeness; not where we differentiate

Python · SQL · pandas / NumPy · scikit-learn / XGBoost · FastAPI · REST / API · Docker · Kubernetes · Git · CI/CD · Airflow · MLflow · Azure · GCP · LangChain

[ 05 ] — WORK

Selected work

Independent Datarmination work — our own methods and tools. Every number below can be recomputed from project artifacts.

R-01
SFIA v9 / ESCO / O*NET / SCOR / BABOK, Python

Professional competency taxonomy

Problem

No coherent, searchable competency structure from job postings.

Approach

A multi-level taxonomy covering thousands of competencies, built via multi-model consensus: different models independently review every disputed assignment, with the outcome decided by majority vote — and every decision recorded.

Value

An auditable hierarchy with a full voting trail and a cross-domain flow map. Low model agreement on disputed cases is proof that a single model is not enough.

R-02
multi-model synthesis, knowledge graph (Graphviz), Python, React

AI glossary 2023–2026

Problem

Fragmented, inconsistent AI terminology with no links between concepts.

Approach

Entries gathered from many independent model sources across successive rounds and linked by a graph of typed relations. In the final round a substantial share of candidates failed a two-stage source verification.

Value

A reproducible reference glossary: source links after automated validation, a maturity rating for every entry, concept evolution chains — fully regenerable from code.

R-03
OpenAI Agents SDK, LangGraph, Pydantic v2 (structured outputs), pytest

Competency classifier — a multi-agent pipeline

Problem

Agentic systems that guess instead of admitting they don’t know.

Approach

Several specialized agents (embedding-based routing, multi-competency splitting, synonyms, taxonomy, translations) + deterministic integrity validation without an LLM. The same system built twice — on OpenAI Agents SDK and LangGraph — with a framework comparison before committing.

Value

Orchestration in code, not by an LLM: closed action codes, confidence thresholds, a human-review path and dozens of tests with a mocked model — an architecture built for production, not for a demo.

R-04
orchestration meta-prompt, several deep-research systems, triangulation

Deep-research orchestration

Problem

A single research model hallucinates and misses sources.

Approach

The same research run in parallel across several independent deep-research systems; a fact counts as certain only once confirmed across independent tracks. Anti-hallucination regime: hard rules, including a mandatory source URL for every claim.

Value

A reusable research template: reports across several thematic tracks, cited sources, documented divergences between tracks.

[ 06 ] — FIRST CALL

How the first call works

  1. 01 30 minutes, no commitments.
  2. 02 We define the problem and the success criterion.
  3. 03 You leave with a scope draft and a package recommendation.

[ 07 ] — FAQ

Frequently asked questions

Why no pricing?+

Because a fair quote depends on your data and scale. We give a fixed price after a 30-minute scoping — upfront, not “as we go”.

What if the project outgrows one person?+

The scope of each stage is clear and finite. I build the core myself; for overload I have a consortium model and trusted partners.

Whose IP is it?+

Yours. The code, prompts and pipelines in full — we deliver them reproducibly, and you run them yourself.

How do we settle up?+

On acceptance of each milestone. You don’t pay for a stage you haven’t accepted.

NDA and data security?+

NDA as standard, before any conversation about data. We work on your infrastructure wherever data can’t leave it.

What happens after the Proof-of-Value?+

Either hardening to production (PoC-to-Production), or an ongoing engagement on a retainer model.

How do you measure LLM response quality?+

A test set + LLM-as-a-Judge + metrics defined upfront. Without that, “it works” is an opinion, not a fact.

[ 08 ] — CONTACT

Contact

Let’s talk about your project — from feasibility analysis to a working demonstrator.

We reply within 24h on business days.