Staff Fullstack Engineer - Internal Tools
LILT
Get hot jobs first on Telegram
New positions appear faster in our channel
- Location
- USA
- Job Type
- full-time
- Work Format
- 🌍 Remote
- Salary
- Remote US $190K – $230K • Offers Equity
- Posted
- October 6, 2026
Job Description
About LILT
AI is changing how the world communicates — and LILT is leading that transformation.
We're on a mission to make the world's information accessible to everyone, regardless of the language they speak. We use cutting-edge AI, machine translation, and human-in-the-loop expertise to translate content faster, more accurately, and more cost-effectively without compromising on brand, voice, or quality.
At LILT, we empower our teammates with leading tools, global collaboration, and growth opportunities to do their best work. Our company virtues—Work together, win together; Find a way or make one; Dance in the customer's shoes; Quicker than they expect; Quality is Job 1—guide everything we do. We are trusted by Intel Corporation, Canva, the United States Department of Defense, the United States Air Force, ASICS, and hundreds of global Enterprises. Backed by Sequoia, Intel Capital, and Redpoint, we’re building a category-defining company in a $50B+ global translation market being redefined by AI.
As part of LILT’s Internal Tools group, you will build and run the tools that LILT's own delivery, contributor-ops, and engineering teams rely on to operate LILT's Applied AI (AAI) benchmarking business — which builds and delivers multilingual benchmarks and evaluation data to frontier AI labs using LILT's global network of subject-matter experts. As the AAI business grows, expect the scope of these platforms to grow in parallel. This platform sits directly on the critical path of paid customer deliverables with real SLAs, maintained by a small, high-leverage engineering team that produces reliable, production-ready tooling. You will own significant surface area across both end-to-end: architecture, data-pipeline design, and the hands-on engineering that keeps this infrastructure reliable. This is a high-impact, high-visibility role: the reliability and craft you bring directly enables on-time, high-quality delivery for some of the most prominent AI labs in the world.
In this role, you will drive the long-term technical strategy for internal platforms while working directly with the delivery, ops, and engineering teams across the business who depend on them daily. This team ships cross-service automation in careful stages — advisory first, then assistive, only later decision-relevant, always with a human fallback — and you'll be expected to hold that same bar as you extend these systems. You will partner with a small existing team to raise the engineering bar across two very different runtimes, set technical direction, and make the calls that determine how this infrastructure scales as the business grows.
What You'll Do
-
Partner with the benchmarking business's researchers and TPMs to translate new benchmark and data-quality requirements into scoped technical designs
-
Own the internal workforce-management tools on our internal platform for AI delivery for hundreds of external contributors: vetting flows, candidate assessment, QC, payment/delivery tracking, and roster/reporting exports
-
Own architecture and long-term technical direction across multiple services and the platform
-
Extend IAA and audio-QA pipelines: annotator outlier detection, ASR sidecar enhancements, LLM-based QC, and DNSMOS/librosa audio-quality scoring
-
Design and ship new modules on the platform that plug new benchmark and vetting workflows into the existing multi-stage review lifecycle
-
Own and extend the platform's API key provisioning and budget-governance system
-
Build self-serve ChatOps-style automation for the internal engineering org and contributor base to enable accelerated annotator workflows and query resolution
-
Harden background worker and job-processing infrastructure
-
Set and enforce testing, CI/CD, and deployment practices across both codebases — from unit/integration testing through infrastructure-as-code and release automation
-
Raise the technical bar for a small, high-leverage team through code review, design docs, and mentoring as the surface area grows
What We're Looking For
Required
-
5+ years of professional full-stack software engineering experience, with a track record of owning production systems end-to-end across more than one runtime/language
-
Deep experience with a Python backend framework (FastAPI or comparable) plus async SQLAlchemy/Postgres, alongside production experience in at least one statically-typed backend language (Go, Java, or similar)
-
Strong React/TypeScript frontend experience — component architecture, state management (Zustand, Redux, or similar), and a modern data-fetching layer (TanStack Query or comparable)
-
Experience building and operating background job/worker systems (queue-driven or polling-based) with failure tolerance and idempotency in mind
-
Experience integrating with third-party and platform APIs — including the GitHub API, OAuth/OIDC SSO, and at least one LLM API (Gemini, OpenAI, or similar) — handling auth, rate limits, and webhook-driven sync
-
Experience building Slack (or comparable chat-platform) bot integrations that automate internal workflows — resource provisioning, approvals, budget/TTL enforcement — with real operational guardrails, not just CRUD features
-
Comfort reading and extending applied-statistics or ML-adjacent code (agreement metrics, audio-quality scoring, or comparable data-quality tooling)
-
Solid grasp of CI/CD, containerized deployment (Docker, Helm, ArgoCD/GitOps or comparable), and infrastructure-as-code (Terraform or comparable)
-
Experience debugging and optimizing native-library (numpy/scipy/onnxruntime-class) memory growth in long-running Python worker processes via safe, boundary-aware process recycling — not just raising memory limits
-
Experience building and owning internal platforms/tools that increase leverage for a non-engineering team (research, operations, support, data/workforce management, or similar) — not solely external-customer-facing product work
Strong Plus
-
Experience with audio/speech pipelines: ASR (Whisper or similar) or audio-quality metrics (DNSMOS, librosa)
-
Experience building internal tools for managing a data-labeling, annotation, or crowdsourced-contributor workforce (vetting, QC, payments)
-
Experience with inter-annotator agreement or statistical agreement metrics
-
Notification and delivery systems experience — Slack bot integrations, transactional email, and idempotent delivery guarantees
-
Experience designing abstractions over heterogeneous data sources with different consistency guarantees — e.g. a fully-replayable event history vs. an observe-only current-state API requiring synthesized diffing — behind one common interface
-
Experience building ChatOps-style automation — Slack or GitHub PR-comment bot commands that trigger backend workflows or CI/CD runs
-
Experience implementing short-lived, rotatable service-to-service JWT auth (key-ID-based rotation, replay-protected tokens, fail-fast config validation) alongside a separate human-facing SSO flow in a paired service
-
Comfort owning both sides of a system with genuinely different runtimes without a large team to lean on
-
Prior experience as the primary or sole engineer on a small, high-leverage internal platform
Bonus
-
Experience with LLM-as-judge or LLM-based QA/review pipelines
-
Familiarity with OpenTelemetry or comparable observability instrumentation in Go services
-
Familiarity with LLM provider gateway/routing services (OpenRouter or comparable) — model aliasing, rate-limit and timeout handling, and budget enforcement
-
Experience with data export/reporting tools (Excel generation, CSV pipelines, or BI-style dashboards)
Our Story
Our founders, Spence and John met at Google working on Google Translate. As researchers at Stanford and Berkeley, they both worked on language technology to make information accessible to everyone. While together at Google, they were amazed to learn that Google Translate wasn’t used for enterprise products and services inside the company.The quality just wasn’t there. So they set out to build something better. LILT was born.
LILT has been a machine learning company since its founding in 2015. At the time, machine translation didn’t meet the quality standard for enterprise translations, so LILT assembled a cutting-edge research team tasked with closing that gap. While meeting customer demand for translation services, LILT has prioritized investments in Large Language Models, human-in-the-loop systems, and now agentic AI.
With AI innovation accelerating and enterprise demand growing, the next phase of LILT’s journey is just beginning.
Our Tech
What sets our platform apart:
-
Brand-aware AI that learns your voice, tone, and terminology to ensure every translation is accurate and consistent
-
Agentic AI workflows that automate the entire translation process from content ingestion to quality review to publishing
-
100+ native integrations with systems like Adobe Experience Manager, Webflow, Salesforce, GitHub, and Google Drive to simplify content translation
-
Human-in-the-loop reviews via our global network of professional linguists, for high-impact content that requires expert review
LILT in the News
-
Featured in The Software Report’s Top 100 Software Companies!
-
LILT makes it onto the Inc. 5000 List.
-
LILT’s continues to be an intellectual powerhouse, holding numerous patents that help power the most efficient and sophisticated AI and language models in the industry.
-
Check out all our news on our website.
Information collected and processed as part of your application process, including any job applications you choose to submit, is subject to LILT's Privacy Policy at .
At LILT, we are committed to a fair, inclusive, and transparent hiring process. As part of our recruitment efforts, we may use artificial intelligence (AI) and automated tools to assist in the evaluation of applications, including résumé screening, assessment scoring, and interview analysis. These tools are designed to support human decision-making and help us identify qualified candidates efficiently and objectively. All final hiring decisions are made by people. If you have any concerns, require accommodations, or would like to opt-out of the use of AI in our hiring process, please let us know at recruiting@lilt.com.
LILT is an equal opportunity employer. We extend equal opportunity to all individuals without regard to an individual’s race, religion, color, national origin, ancestry, sex, sexual orientation, gender identity, age, physical or mental disability, medical condition, genetic characteristics, veteran or marital status, pregnancy, or any other classification protected by applicable local, state or federal laws. We are committed to the principles of fair employment and the elimination of all discriminatory practices.
🎯 Who is this job for?
This role is suitable for a Senior to Staff-level full-stack engineer with 5+ years of experience owning production systems end to end across Python and a statically typed backend language, plus React/TypeScript expertise.
The person should be skilled in API and platform architecture, PostgreSQL, asynchronous workers, third-party and LLM integrations, ChatOps automation, CI/CD, Docker, GitOps, infrastructure as code, observability, and reliable process management for long-running Python services.
They should be familiar with designing and operating internal tools for research, operations, and contributor workforces, including vetting, quality control, payments, reporting, data and audio pipelines, workflow automation, testing, deployment, and mentoring a small high-leverage engineering team.
💬 Potential Interview Questions
How would you design a FastAPI service using async SQLAlchemy and Postgres for a multi-stage contributor vetting workflow?
I would model each workflow stage and transition explicitly, enforce state-transition rules transactionally, and use async SQLAlchemy with connection pooling. I would add audit events, idempotency keys, indexes for operational queries, and integration tests covering concurrent updates and retries.
How would you make background workers reliable when processing annotation, QC, or delivery jobs?
I would use durable queues, explicit job states, visibility timeouts or leases, bounded retries with backoff, and a dead-letter path. Each handler would be idempotent, persist progress transactionally, and emit metrics and structured logs for latency, failures, and retry volume.
How would you prevent memory growth in long-running Python workers using NumPy, SciPy, or ONNX Runtime?
I would measure process RSS and library-specific behavior, isolate workloads in child processes, and recycle workers after a bounded number of jobs or when memory thresholds are exceeded. This is safer than only increasing limits because native allocations and fragmentation may not be released reliably within the Python process.
How would you design API key provisioning and budget governance for internal services and LLM providers?
I would store secrets in a dedicated secret manager, expose short-lived or scoped credentials where possible, and maintain ownership, limits, expiration, and usage records. Provisioning would require authorization, enforce budgets and TTLs, support rotation and revocation, and fail closed when governance data is unavailable.
What considerations are important when integrating GitHub, OAuth/OIDC SSO, Slack, and LLM APIs?
I would handle token encryption and rotation, validate OAuth state and OIDC claims, verify webhook signatures, and design for provider rate limits, pagination, retries, and partial failures. External calls should be time-bounded, observable, and reconciled through idempotent synchronization jobs rather than assumed to be immediately consistent.
How would you implement a Slack ChatOps workflow for approvals, resource provisioning, and budget enforcement?
I would validate the requesting user and command context, authorize actions against current policy, and require confirmation for consequential operations. The backend would create an idempotent workflow with audit records, expiration and budget checks, clear status updates, and a safe fallback when Slack or downstream services are unavailable.
How would you structure a React and TypeScript frontend for complex internal workflows?
I would separate reusable presentational components from workflow-specific containers, use typed API contracts, and manage server state with TanStack Query while reserving Zustand or a similar store for client state. I would include optimistic updates only where safe, handle loading and error states explicitly, and test critical user journeys.
How would you design an abstraction over replayable event histories and observe-only current-state APIs?
I would define a common domain interface that returns normalized entities, versions, timestamps, and synchronization metadata while keeping source-specific adapters separate. Replayable sources could derive state directly, whereas current-state sources would require snapshots, synthesized diffs, and reconciliation rules that clearly document consistency limitations.
How would you build a multilingual audio-quality or ASR quality pipeline using tools such as librosa, DNSMOS, and Whisper?
I would normalize and validate audio inputs, isolate CPU- and memory-intensive processing, and persist intermediate artifacts and versioned scores for reproducibility. Quality checks would include sample-rate and duration validation, model and metric versioning, threshold calibration, and human review for low-confidence or anomalous results.
What CI/CD and deployment practices would you establish across Python and statically typed services?
I would standardize formatting, linting, type checking, unit and integration tests, container builds, security scanning, and migration checks in CI. Deployments would use immutable images, Terraform-managed infrastructure, Helm and GitOps-style promotion, automated rollback signals, and observability through logs, metrics, traces, and service-level indicators.
📋 Job Summary
LILT is an AI-powered language technology company making global information accessible through machine translation and human-in-the-loop expertise. As a Staff Fullstack Engineer on the Internal Tools team, you’ll own production platforms for workforce management, benchmark delivery, data-quality pipelines, ChatOps automation, APIs, background workers, and reliability. The stack includes Python, FastAPI, async SQLAlchemy, PostgreSQL, Go, React, TypeScript, Docker, Helm, ArgoCD, Terraform, Slack, and LLM APIs. This is a fully remote US role offering $190K–$230K plus equity—an excellent opportunity to shape high-impact infrastructure used to deliver critical AI benchmarks for leading global labs.
Required Skills
Never miss a JavaScript opportunity
Subscribe to get similar jobs and weekly insights delivered to your inbox
Hiring JavaScript developers?
Post your job to 9,100+ registered developers. Starting free.
See PricingRelated jobs
Is this your listing? Claim or request removal