Staff Software Engineer, Observability & Profiling

Anthropic

Staff Software Engineer, Observability & Profiling Overview

Company Name Anthropic
Job Role Staff Software Engineer, Observability & Profiling
Qualifications Not Specified
Category IT Jobs
Job Type Full Time
Location London

Anthropic is seeking experienced Software Engineers to join their Observability team within the Infrastructure division. This team is responsible for the monitoring and telemetry infrastructure that supports all engineers and researchers at Anthropic. The systems they build include metrics and logging pipelines, distributed tracing, profiling, error analytics, alerting mechanisms, and dashboards that enable quick and effective troubleshooting and system understanding. As part of this team, you will directly influence the reliability and operational excellence of Anthropicâ??s research and product systems.

About the Role

In this position, you will focus on designing and constructing scalable telemetry ingestion and storage solutions for metrics, logs, traces, and error data across multiple clusters. You will develop observability tools that provide engineers with detailed, low-overhead insights into system behavior across the entire infrastructure. Your responsibilities include maintaining and improving core observability platforms, leading migration efforts, and implementing architectural enhancements to improve system reliability, reduce costs, and support growth.

You will also create instrumentation libraries, SDKs, and leverage eBPF-based auto-instrumentation techniques to generate high-quality telemetry data with minimal code changes. Your work will help reduce the mean time to detect and resolve issues by enabling cross-signal correlationâ??from kernel events to application tracesâ??and building unified query interfaces, including AI-assisted diagnostic tools. Additionally, you will turn continuous profiling and resource utilization telemetry into actionable insights to optimize CPU, memory, and accelerator performance across the fleet.

Key Responsibilities

  • Design and implement scalable telemetry pipelines for metrics, logs, traces, and errors across Anthropicâ??s multi-cluster environment.
  • Develop observability solutions that offer deep, low-overhead visibility into system operations for engineering teams.
  • Own and evolve core observability platforms, leading migration efforts and architectural improvements to enhance reliability, reduce costs, and support organizational scaling.
  • Create instrumentation libraries, SDKs, and utilize eBPF-based auto-instrumentation to emit high-quality telemetry data with minimal disruption.
  • Reduce detection and resolution times by building cross-signal correlation, unified query interfaces, and integrating AI-powered diagnostic tools.
  • Drive fleet-wide efficiency by transforming profiling and telemetry data into actionable optimization insights for CPU, memory, and accelerator resources.
  • Collaborate with Research, Inference, Product, and Infrastructure teams to ensure observability solutions are tailored to their specific operational needs.

Qualifications

  • Practical experience in constructing and managing large-scale observability or monitoring systems.
  • Deep understanding of the entire signal pipelineâ??from instrumentation to data ingestion, querying, and analysis.
  • Knowledge of high-throughput telemetry systems and the tradeoffs involved in data collection, storage, and querying at scale.
  • Ability to investigate below the application layer, including kernel, network stack, or hardware components.
  • Excellent communication skills and a collaborative approach to working with internal teams on operational visibility and incident response.
  • Motivated to build foundational infrastructure and capable of tackling complex technical challenges independently and with a team.

Preferred Qualifications

  • Over 10 years of relevant industry experience, including managing large-scale observability systems.
  • Experience with eBPF-based observability in production environments, including tracing, profiling, or network visibility.
  • Experience running continuous profiling at fleet scale, managing overhead budgets, and symbolization.
  • Kernel and syscall-level debugging skills and performance engineering expertise.
  • Experience profiling or instrumenting workloads on accelerators.
  • Experience operating metrics systems with high cardinality or large telemetry storage backends.
  • Familiarity with OpenTelemetry instrumentation, collector pipelines, and tail-based sampling strategies.
  • Interest in applying AI and large language models to operational workflows such as root cause analysis, anomaly detection, or alerting.

Compensation and Benefits

The annual salary for this role ranges from £325,000 to £390,000 GBP. The company offers flexible working arrangements, with a hybrid policy requiring staff to be in the office at least 25% of the time, though some roles may require more in-office presence. They sponsor visas and will make reasonable efforts to secure one for successful candidates, employing legal support for immigration processes. Additional benefits include competitive pay, optional equity donation matching, generous vacation and parental leave, and a well-designed office space for collaboration.

Additional Information

Anthropic values diversity and encourages candidates from underrepresented groups to apply, emphasizing that not all qualifications are mandatory. They prioritize impact-driven research and foster a collaborative environment with frequent research discussions. The company is headquartered in San Francisco and is committed to social and ethical considerations in AI development. Candidates are advised to apply via official channels and be cautious of scams, with legitimate recruiters contacting only from @anthropic.com email addresses.


Degree Requirement: Not Specified

Visa Sponsorship Promising

To apply for this job please visit job-boards.greenhouse.io.