We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results

Senior Site Reliability Engineer

Spectraforce Technologies
United States, Texas, Austin
2435 East Riverside Drive (Show on map)
Sep 10, 2026
Title: Senior Site Reliability Engineer

Duration: 6 Months (Could extend upto 18 months)

Location: Austin, TX - Hybrid 4 days weekly onsite

Qualifying

Reason for Opening? support for moving from on-prem to cloud, infra support etc

Why is this role important to your team/the project/the company? We are looking for a skilled engineer with disciplines that incorporate aspects of software systems engineering and operations. We are combining these skills to come up with better ways of managing and operating applications - including AI/ML-driven approaches to observability and reliability.

Our Opportunity:

We are looking for a skilled engineer with disciplines that incorporate aspects of software systems engineering and operations. We are combining these skills to come up with better ways of managing and operating applications - including AI/ML-driven approaches to observability and reliability.

What you'll do:

* Evangelize SRE mindset and solve problems through systematization.

* Identify opportunities to build innovative tools and solve unique operations problems on large enterprise and mission-critical applications.

* Create scripts to automate operational tasks and incorporate solutions into infrastructure; architect and own production automation solutions that measurably reduce manual toil and improve operational throughput.

* Design and implement AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems - including anomaly detection and predictive alerting to improve platform reliability.

* Lead expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms.

* Collaborate with Engineering, Scrum, and Ops resources to provide technical expertise and support on key initiatives for system availability and reliability.

* Triage alerts and diagnose/resolve critical issues; manage implementation of changes with clear communication and minimal risk.

* Develop tools, frameworks, and instrumentation to validate and increase rollout success for applications; leverage AI/ML capabilities to enhance operational visibility and rollout validation at scale.

* Champion AIOps platform adoption and ML-assisted observability practices across the team.

* Coordinate capacity planning using data-driven trend analysis and ML-informed forecasting.

* Develop CI/CD orchestration systems to reduce friction for software delivery to production; drive adoption of GitOps concepts and AI-assisted pipeline optimization.

* Real-time troubleshooting of mission-critical application workflows and incorporate feedback into product development.

* Participate in on-call support.

What do you have:

Required Skills:

* 6-8 years of experience with enterprise-level administration and support.

* 6-8 years of experience writing automation scripts, building application dashboards for proactive monitoring, and setting up alerts for early issue determination.

* 6-8 years practicing SDLC, process improvements.

* Hands-on enterprise systems administration, monitoring, and deployment activities.

* Experience with Windows 2019/2022 and Linux hosted via Virtual Machine.

* Experience in Cloud application configuration, deployment, support, and migration - GCP/PCF is a plus.

* Knowledge of IP networking including DNS, DHCP, firewalls, IP routing, etc.

* Familiarity with large-scale distributed systems and high-availability architecture.

* Linux and Windows system administration, troubleshooting, and tuning.

* Development experience in one or more programming languages: .NET, PowerShell, Java, Python, Bash.

* Knowledge of one or more of SQL, Oracle, MongoDB databases.

* Working knowledge of Actimize.

* Knowledge of one or more Message Brokers: Solace, RabbitMQ, IBM MQ, Kafka.

* Knowledge of Splunk, AppDynamics, or similar observability tools.

* Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML-assisted observability) in production environments.

* Bachelor's degree in computer science or related discipline.

Helpful Skills:

* Financial services industry experience.

* Agile methodologies.

* Hands-on experience with AIOps platforms or ML-driven observability tooling.

* Experience integrating AI/ML capabilities into CI/CD or operational automation workflows.

* Familiarity with CI/CD tools (Harness, Jenkins, GitHub Actions) or GitOps concepts.

* Exposure to container orchestration (Kubernetes, OpenShift) or cloud platforms (AWS, Azure, GCP).

Personal Skills:

* Strong customer orientation with an affinity to proactively own, communicate, and follow through on projects and issues.

* Extreme sense of ownership to resolve problems in a distributed environment.

* Gritty resolve to dig deeper into technical issues in a complex login ecosystem.

* A self-starter with the ability and confidence to independently resolve issues and bring results back to the team.

Must Have

  • AI tools - Claud, co pilot (none specific)

    • Must have strong understanding of AI ops and hands on exp


  • Tools
  • GCP or other cloud platforms (more than 1)
  • Kubernetes
  • Terraform
  • Python


Notes from my Call:

This is an IC role

Responsible for virtual application resiliency and sustainability 'expected to be working on all Ai assisted automation and AI observability building , defend system and production and platform issues

Handel cross platform communication

Cloud - GCp exp or at least exp with 2 other cloud platforms (azure/Aws etc)

Test , setup up and take it into production

And AI hands on

SRE proactive and knowledge

4 main areas they will work on:

  • production issue is primary
  • Observability
  • infra set up
  • work on infra set up and building that infra
  • Moving to the could 'Kubernetes terraform and moving to GCP cloud
  • Moving from on prem to cloud
  • Team owns setting up the infra
  • Anything/everything manual they aim to automate it


Project:

Critical application SRE support team project

Work ranges from - production availably stability and resiliency

Always monitoring 24/7 monitoring

Take action on alerts

Make sure application is stable

be proactive not reactive

Monitor signals and take action before it happens

Reposting to originality is priority

Monitoring is AI driven and continuously and issues are caught before issue happens

Observability of application

Production issues

AI - 50/50:

As and when needed jump in a resolve production issues

Remining time will be infra and automation

QUESTIONS FROM TEAM:

  • Could you please clarify how critical AIOps platform experience is for this role? Is it a must-have requirement, or would you consider it nice to have if the candidate has strong experience in the core areas of the role?

    • yes this is must have skill





Interview Process:

1 st - open book assessment (complete asap but have 1 week MAX to complete)

2 nd - on-site technical intv (can be virtual if not local) - 1 hr

3 rd - final 30 min manager intv

Applied = 0

(web-665cd84569-2d8ll)