tialuk · · 5 min read

tialuk: selected code

Five short excerpts from a private codebase: security, SQL, AI privacy, integrations and deploys.

tialuk's repository is private, so instead of the whole thing here are five excerpts, lightly simplified, that show how I approach different kinds of problems. None of them is enough to rebuild the product, and I'm happy to walk through the rest in an interview.

Constant-time origin check

Only traffic that went through the Cloudflare edge should reach the backend, so the edge injects a shared secret header and this middleware rejects everything else, comparing in constant time so response latency can't leak the secret. A missing secret in production is a boot failure instead of a silent passthrough, and /health stays exempt because the Cloud Run startup probe never carries the header.

import { Elysia } from "elysia";
import { timingSafeEqual } from "node:crypto";

const EXEMPT_PATHS = new Set(["/health", "/healthz"]);

export function originGuard() {
  const rawSecret = process.env.EDGE_SHARED_SECRET;
  if (!rawSecret) {
    if (process.env.NODE_ENV === "production") {
      throw new Error("EDGE_SHARED_SECRET is required in production: refusing to boot without the origin guard.");
    }
    return new Elysia({ name: "origin-guard" }); // dev: passthrough, warned once at boot
  }
  const secretBuf = Buffer.from(rawSecret);

  return new Elysia({ name: "origin-guard" }).onBeforeHandle({ as: "global" }, ({ request, set }) => {
    if (EXEMPT_PATHS.has(new URL(request.url).pathname)) return;
    const received = request.headers.get("x-edge-auth");
    let ok = false;
    if (received !== null) {
      const receivedBuf = Buffer.from(received);
      // timingSafeEqual throws on different lengths, so check the length first
      if (receivedBuf.length === secretBuf.length) ok = timingSafeEqual(receivedBuf, secretBuf);
    }
    if (ok) return;

    set.status = 403; // logged at most once per second, never with the full header value
    return { error: "forbidden" };
  });
}

Purging the ledger without moving the balance

A tenant's AI credit balance is the sum of its credit transactions minus the cost of the usage rows still in the ledger, so a plain retention purge would quietly push every balance up. Deleting the old rows and writing one consolidation entry per tenant in the same statement keeps the total identical, is atomic by definition and needs no explicit locking across replicas: the replica that loses the race deletes nothing and consolidates nothing.

WITH del AS (
  DELETE FROM ai_usage_ledger WHERE created_at < ? RETURNING tenant_id, credit_cost
), ins AS (
  INSERT INTO credit_transactions (tenant_id, type, amount, note, metadata, created_at)
  SELECT tenant_id, 'consolidation', -SUM(credit_cost),
         'retention sweep consolidation',
         jsonb_build_object('consolidatedRows', COUNT(*), 'cutoff', ?::text),
         now()
    FROM del GROUP BY tenant_id HAVING SUM(credit_cost) > 0 -- failed or pre-credit rows cost 0
  RETURNING id
)
SELECT COUNT(*)::int AS deleted FROM del;

Redacting PII before the model sees it

Customer messages go to an LLM, so emails, IBANs, card numbers, national IDs and phone numbers are swapped for numbered placeholders that keep the meaning ("the customer's email") without the data. The pass order is the real design: specific patterns run before broad ones (a 16-digit card also matches the phone regex), and card candidates must pass a Luhn check so a random digit run falls through to the later passes.

const PASSES: PiiPass[] = [
  { kind: "EMAIL", pattern: EMAIL_PATTERN },
  { kind: "IBAN", pattern: IBAN_PATTERN },
  { kind: "CARD", pattern: CREDIT_CARD_PATTERN, validate: luhnCheck },
  { kind: "CUIT", pattern: CUIT_AR_PATTERN }, // 11 digits, before the 7-8 digit DNI
  { kind: "DNI", pattern: DNI_AR_PATTERN },
  { kind: "PHONE", pattern: PHONE_PATTERN },
];

export function redactPIIForAI(text: string): RedactResult {
  if (!text) return { redacted: text ?? "", replacements: [] };
  let working = text;
  const replacements: PIIReplacement[] = [];
  for (const pass of PASSES) {
    const matches: Array<{ start: number; end: number; value: string }> = [];
    for (const m of working.matchAll(pass.pattern)) {
      if (m.index === undefined) continue;
      if (pass.validate && !pass.validate(m[0])) continue;
      matches.push({ start: m.index, end: m.index + m[0].length, value: m[0] });
    }
    // Replace right to left so earlier indices stay valid; the leftmost match is still _1.
    for (let i = matches.length - 1; i >= 0; i--) {
      const { start, end, value } = matches[i]!;
      const placeholder = `[${pass.kind}_${i + 1}]`;
      working = working.slice(0, start) + placeholder + working.slice(end);
      replacements.push({ kind: pass.kind, placeholder, originalLength: value.length });
    }
  }
  return { redacted: working, replacements }; // (audit metadata is sorted first, omitted)
}

Self-healing integration tokens

A channel's long-lived access token can die silently, and the customer only notices when messages stop arriving. The periodic health check separates transient failures (rate limits, network), which just reschedule, from terminal ones, which count strikes; on the third strike it tries to re-derive the token from the stored user token, and only when that fails does it flag the channel for a manual reconnect.

const err = result.error; // the token probe failed

// Transient: the token may be fine, the connection to Meta is not. Reschedule, no strike.
const isRetryable =
  err instanceof MetaRateLimitError ||
  (err instanceof MetaApiError && err.code === -1);
if (isRetryable) {
  await deps.recordFailure(channel.id, {
    status: channel.healthStatus,
    error: err.message,
    nextCheckAt: new Date(Date.now() + RETRY_AFTER_TRANSIENT_MS), // 30 min
    incrementFailures: false,
  });
  return;
}

// Terminal (expired token, invalid param): strike 1 stays HEALTHY, strike 2 is DEGRADED.
const nextFailures = channel.consecutiveHealthFailures + 1;
if (nextFailures < 3) {
  await deps.recordFailure(channel.id, {
    status: nextFailures === 1 ? HealthStatus.HEALTHY : HealthStatus.DEGRADED,
    error: err.message,
    nextCheckAt: new Date(Date.now() + RETRY_AFTER_FIRST_FAIL_MS), // 1 h
    incrementFailures: true,
  });
  return;
}

// Strike 3: re-derive from the user token. Success resets to HEALTHY, failure is NEEDS_RECONNECTION.
await attemptReDerive(channel, err, deps);

A release gate for destructive migrations

Production deploys promote the exact image that already ran in dev, and the automatic rollback only moves traffic back to the previous revision: it can't undo a schema change. So before the manual approval, the pipeline lists the migrations the release adds and blocks any drop, truncate or column type change unless the release notes carry an explicit, auditable [allow-destructive] opt-in.

- name: Migration plan
  env:
    RELEASE_BODY: ${{ github.event.release.body }}
  run: |
    # prev_sha: commit of the image prod is running now; release_sha: the release tag.
    new_migrations=$(git diff --name-only --diff-filter=A \
      "$prev_sha".."$release_sha" -- 'db/migrations/Migration*.ts')
    [ -z "$new_migrations" ] && exit 0

    # Additive changes (e.g. ALTER TYPE ... ADD VALUE on enums) don't match.
    destructive=""
    for f in $new_migrations; do
      if git show "$release_sha:$f" | grep -Eiq 'drop table|drop column|truncate|alter column[^;]*type'; then
        destructive="$destructive $f"
      fi
    done

    if [ -n "$destructive" ]; then
      case "$RELEASE_BODY" in
        *"[allow-destructive]"*)
          echo "Release body contains [allow-destructive]: explicit opt-in, continuing."
          ;;
        *)
          echo "::error::Destructive migrations:$destructive. Add [allow-destructive] to the release body (rollback does NOT revert them)."
          exit 1
          ;;
      esac
    fi

← Back to the portfolio · hello@abrahamkazerian.dev