Tech Lead at Growtrics / Engineering practice

Build the system.
Own the decisions.

Architecture, performance, delivery, and the operational work that connects them. This is what I’ve built, what I’ve measured, and where the experiments stopped.

ProductexperienceRuntimeReleasePublishingLifecycleState & contractsArtifacts & policyContent & authorityIdentity & context
The work between a feature and a dependable product.

Across the system.

Responsibility beyond individual features

01TaloTrace / platform architecture

One platform, explicit service ownership.

Agent runtimes, durable workflows, typed contracts, and the boundaries between them.

Applied stack
  • Python
  • Agno
  • PostgreSQL
  • Typed contracts
  • Cloud Run
01Conceptual system map
Application ownsWorkflow state
& execution budgets
Durable processes
Typed contracts
Framework executesNative Agno composition
AgentModelTools
Shared engineering guidance across 14 repositories
Responsibility map: workflow state and budgets belong to the application; agent execution belongs to the framework.
Read the decision & outcome

The problem

Exploration, navigation, evidence, and orchestration have different responsibilities. Shared execution logic can easily become another layer of duplicated state.

My decision & contribution

I worked across service boundaries and shared contracts, introduced durable process primitives, and refined native Agno composition. The application owns workflow state and budgets; framework components handle agent execution.

The result

The work established reusable execution foundations and shared coding/review guidance merged across 14 TaloTrace repositories. The native composition changes reached development, with a recorded 3,005 local tests and deployment smoke checks.

Development integration and checks are recorded. A framework migration alone does not establish faster or more reliable agent behavior.

02Growtrics & TaloTrace / release engineering

A release is an artifact and a policy.

Mobile store delivery, OTA updates, shared CI, promotion, and rollback.

Applied stack
  • Flutter
  • Fastlane
  • Shorebird
  • GitHub Actions
  • Cloud Build
  • GCP
02Conceptual system map
Traceable inputsSource & artifacts
Store buildSigned packageFastlane
OTA patchCompatible updateShorebird
BackendImage promotionCloud Build
ValidationPromotionRecovery
Delivery map: store builds, OTA patches, and backend promotions retain their own validation and recovery rules.
Read the decision & outcome

The problem

A signed app-store build, an OTA patch, and a backend promotion have different compatibility rules. Rebuilding or changing inputs between environments makes the release harder to trace.

My decision & contribution

I codified store, Shorebird release, and patch delivery modes, then worked on immutable source/image promotion and dependency-ordered rollout. Shared conventions coexist with repository-owned CI and explicit recovery paths.

The result

The release tooling makes version inputs, validation, and delivery modes inspectable. A historical Android production-flavor release was verified through Shorebird and Google Play internal distribution.

Internal store distribution is distinct from a public app-store rollout. The delivery modes have different validation and compatibility requirements.

03Growtrics / content and publishing architecture

Give editors ownership without losing control.

CMS migrations, structured content, live preview, media delivery, and publication authority.

Applied stack
  • Next.js
  • Payload CMS
  • Sanity
  • PostgreSQL
  • Mux
  • Vercel Blob
  • MCP
03Conceptual system map
Editor changesBounded agent edits
Content workspaceDraft → Preview
Human publication authority
Published contentLayout stays in the application
Publication map: editors and bounded agent edits enter a controlled content workflow. Humans retain publication authority.
Read the decision & outcome

The problem

Content in application source ties editorial changes to engineering releases. Moving it into a CMS also changes schemas, rendering, media, and who can publish.

My decision & contribution

I migrated the earlier content backbone from TinaCMS to Sanity, then delivered the newer company site’s Payload architecture. I separated layout from editorial data and added controlled migrations, drafts, live preview, and bounded agent editing.

The result

The company site’s CMS is recorded in production. Editors can manage content without an application deploy, and the recorded migration check preserved the public text across 24 pages. Agent writes remain drafts under human publication authority.

My contribution is the CMS and operating architecture. The site’s visual design was created separately. The engagement-reporting interface exists, but populated live analytics was not verified.

04Growtrics / CRM and customer operations

Connect acquisition to the product lifecycle.

Lead intake, attribution, identity, trial state, and lifecycle communication.

Applied stack
  • HubSpot
  • Customer.io
  • Next.js
  • FastAPI
  • Webhooks
  • Product events
04Conceptual system map
ContactWaitlistRegistration
Carry across the boundaryIdentity + attribution
HubSpotAcquisition flows
Customer.ioLifecycle communication
AccountTrialCampaign
Integration map: acquisition context and product lifecycle state cross service boundaries without losing identity.
Read the decision & outcome

The problem

Landing-page submissions and account state live in different systems. A form integration is incomplete if attribution or customer identity disappears across that boundary.

My decision & contribution

I connected historical acquisition flows to HubSpot and product lifecycle communication to Customer.io. The work carries page context through registration and synchronizes account, trial, and campaign state.

The result

The integrated paths cover contact, waitlist, and registration intake, alongside transactional communication and lifecycle webhooks. I also documented the tracking map for marketing to inspect.

This is CRM integration and lifecycle engineering. Conversion uplift and campaign effectiveness were not measured. The current company-site contact demo is a separate implementation.

Measured, with context.

Recorded checks, with their boundaries attached

These results come from recorded engineering checks. Each comparison keeps its method and scope attached.

01Recorded experiment

Remote command latency

Read-only browser commands across cloud placements.

TaloTrace / recorded placement experiment

Before223.3 ms
After30.5 ms
86.4% lower median read latency in the recorded probe.

Method & conditions

6 October 2026. Three fresh browser allocations per region, with 20 read commands per allocation. Compare pooled command medians.

What this establishes

Two read-only browser methods. Excludes startup, lease acquisition, connection setup, model calls, and the full QA journey. This is a placement experiment, not a production-wide speedup.

02Recorded experiment

Two-field action groups

Warm local two-field execution, without model inference.

TaloTrace / controlled local execution

Before1.203 s
After0.327 s
72.8% lower median execution time with ordered calls.

Method & conditions

6 October 2026. Eight two-field groups per variant. The local actor records observations, screenshots, and checkpoint writes.

What this establishes

Warm local execution. Excludes model inference, remote storage, SQL, and live publication. The result does not establish full-agent latency or a production rollout.

03Migration verification

Public content preserved

Recorded production CMS migration handoff.

Growtrics / recorded live migration check

Public URLs24 pages
Text parity24 matched

One mark for each public page checked

Zero text differences after the CMS migration.

Method & conditions

9 October 2026 deployment handoff. Recorded parity across 24 public URLs, plus desktop and mobile interaction checks.

What this establishes

This is migration correctness, not a speed benchmark. The recorded handoff was inspected; the live parity run was not repeated for this portfolio.

04Recorded experiment

Reasoning effort in a lab journey

3 high-effort runs and 8 medium-effort runs; all passed the lab checks.

TaloTrace / recorded reasoning-effort experiment

High effort90.5 s
Medium effort65.7 s
Observed median at high vs medium reasoning effort.

Method & conditions

7 October 2026. Three high-effort runs and eight medium-effort runs of the same login-form journey, engine version and model. Both used a real cloud browser, local PostgreSQL with 20 ms added delay and a 20 ms orchestration stub. All 3 high and 8 medium runs passed the lab checks.

What this establishes

Small, non-interleaved samples with variable browser waits; this does not establish unchanged general accuracy or a production speedup. Timing includes browser lease, launch, model calls and recording; final verification readback is excluded.

The experiment log.

Including the paths I did not ship

  1. 01Compared

    Isolate the loop cost before replacing the browser library

    I compared execution paths using equivalent browser-command protocols. The library comparison alone did not explain the delays. The investigation instead separated network round trips, repeated observations, model turns, and readiness waits.

    Browser Use · CDP · action-group experiments

    Scope of the result

    A faster executor can still sit inside a slower agent loop. End-to-end runs with different decisions do not establish a causal speedup.

  2. 02Development deployed

    Use native framework components, then test the behavior

    I compared the existing agent design with native framework composition and revised the integration around real agent, model, tool, and typed-output components. Shared runtime helpers stayed focused on execution checks.

    Agno · Python · typed outputs · development smoke tests

    Scope of the result

    Local tests and development deployments are recorded. A change in framework or model cost is not, on its own, a demonstrated behavior or latency improvement.

  3. 03Historical experiment

    Explore GPU-backed speech serving

    I explored GPU-backed speech generation using Modal, vLLM, and Higgs Audio. A short-lived shared prototype was removed; a separate historical speech-serving implementation followed. The work covers serving integration, rather than authorship of the underlying model.

    Modal · vLLM · PyTorch · HiggsAudio

    Scope of the result

    Historical implementation evidence exists. Current deployment, throughput and voice-latency gains were not verified.

  4. 04Prototype

    Prototype faster projection writes while preserving recovery

    I evaluated batched Redis writes against a Lua-based atomic path, with recovery checks alongside latency measurements. The prototype remained outside the merged delivery path.

    Redis · Lua · PostgreSQL · recovery tests

    Scope of the result

    The experiment used a local database and remote Redis with synthetic notification delay. It is not a customer-performance result.

  5. 05Development acceptance

    Build a device fleet around actual demand

    I worked on per-device allocation and scale-to-zero instead of keeping a large shared host warm. The recorded development acceptance tested ten simultaneous leases, repeated acquisition, the capacity ceiling, and cleanup back to zero.

    GCP · Android execution · idempotent leases · crash recovery

    Scope of the result

    The session reports development acceptance. Extended workload stability and production promotion remained separate gates. Cost comparisons were models, not settled invoice savings.

  6. 06Investigated

    Make benchmark integrity a prerequisite for a score

    I investigated navigation benchmarks where dataset preconditions shifted step evidence and invalid judge responses affected eligibility. The work separated harness failures from agent failures and preserved historical scorecards while defining a corrected verification cohort.

    Navigation evaluation · evidence provenance · screenshot judging

    Scope of the result

    Invalid or incomparable runs are not treated as improvements in agent accuracy.

  7. 07Lab comparison

    Test a model change before changing the product

    I tested an alternative model with explicit prompt caching against the current exploration policy. In one three-run comparison, the candidate’s median journey took 44.8 seconds versus 41.3 seconds for the baseline; all six runs completed correctly. The small sample did not justify a model switch.

    Python · OpenRouter · prompt caching · browser-agent evaluation

    Scope of the result

    October 2026 lab runs on one login-form journey. Provider timing and agent decisions vary; this is a recorded comparison, not a general ranking of models. Cost observations came from separate instrumented runs.

  8. 08Exploratory lab comparison

    Separate infrastructure delay from agent behavior

    I built an experiment that changed database placement and simulated orchestration delay around the same agent version. One run took 124.8 seconds with the delayed setup; one took 76.7 seconds with a regional database and no added orchestration delay. Both passed the lab checks.

    PostgreSQL · Toxiproxy · cloud browsers · Python

    Scope of the result

    One run per setup. Database, orchestration delay and browser waits differed, so this observation does not isolate a database-only gain or establish a production speedup.

See the product case studies