We are looking for a NOC Engineer responsible for monitoring, managing, troubleshooting, and maintaining the availability and performance of our network, servers, telephony, and production infrastructure.

The candidate will work closely with the DevOps, SIP/VoIP, CTI, Development, and QA teams to identify incidents, troubleshoot infrastructure and connectivity issues, and ensure timely resolution of production-impacting problems.

Key Responsibilities
  1. Monitor production infrastructure, servers, networks, applications, and telephony services 24×7 as per the defined shift schedule.
  2. Monitor CPU, memory, disk, network utilization, latency, packet loss, service availability, and system health.
  3. Monitor SIP/VoIP infrastructure, SIP trunks, channels, call connectivity, registration, call failures, and related telephony services.
  4. Troubleshoot network connectivity issues involving TCP/IP, DNS, HTTP/HTTPS, VPN, firewalls, routing, and ports.
  5. Monitor AWS infrastructure, including EC2, ECS, RDS, Load Balancers, CloudWatch, and other relevant services.
  6. Identify production incidents proactively through monitoring and alerting systems.
  7. Perform first-level troubleshooting and determine whether an issue is related to network, server, application, database, SIP/telephony, or third-party services.
  8. Escalate incidents to the appropriate DevOps, Development, SIP/VoIP, Database, or third-party team with complete technical details.
  9. Maintain incident logs, escalation records, and resolution documentation.
  10. Perform health checks before and after production deployments.
  11. Monitor service uptime, system performance, and infrastructure capacity.
  12. Coordinate with telecom/SIP providers for issues related to SIP trunks, channels, numbers, call routing, connectivity, and call quality.
  13. Analyze logs and network-level information to identify the root cause of incidents.
  14. Support troubleshooting using tools such as Ping, Traceroute, Telnet/Netcat, nslookup/dig, curl, SSH, TCPDump, Wireshark, and similar utilities.
  15. Monitor alerts and ensure incidents are acknowledged and resolved within defined SLA/response timelines.
  16. Prepare daily/weekly operational reports covering incidents, outages, recurring issues, and system health.
  17. Maintain proper documentation for infrastructure, monitoring, escalation matrices, and standard operating procedures.
  18. Participate in incident management, problem management, and post-incident RCA activities.
  19. Support infrastructure capacity planning by identifying trends in resource utilization and service demand.

Required Technical Skills
  1. Networking
    1. Strong understanding of TCP/IP, UDP, DNS, DHCP, HTTP/HTTPS, VPN, NAT, routing, and firewalls.
    2. Understanding of network troubleshooting and connectivity diagnostics.
    3. Knowledge of LAN/WAN and network monitoring.
    4. Basic understanding of load balancing and reverse proxies.
  2. Linux & Servers
    1. Good working knowledge of Linux/Ubuntu/CentOS environments.
    2. Comfortable with SSH and Linux command-line troubleshooting.
    3. Understanding of processes, services, disk, memory, CPU, and network troubleshooting.
    4. Basic shell scripting knowledge is an advantage.
  3. AWS / Cloud
    1. Working knowledge of AWS infrastructure.
    2. Familiarity with EC2, ECS, RDS, CloudWatch, IAM, VPC, Security Groups, Load Balancers, and related services.
    3. Ability to troubleshoot basic AWS connectivity and infrastructure issues.

SIP / VoIP / Telephony
  1. Experience with SIP/VoIP or telecom environments will be highly preferred.
  2. The candidate should have an understanding of
    1. SIP signaling and call flow
    2. SIP registration
    3. INVITE / BYE / ACK / CANCEL / OPTIONS
    4. SIP response codes
    5. RTP and media flow
    6. SIP trunking
    7. Concurrent channels
    8. Call routing
    9. Caller ID
    10. Call failures and disconnects
    11. Basic VoIP call-quality parameters such as latency, jitter, and packet loss

Monitoring & Troubleshooting
  1. Experience with monitoring/logging tools such as -
  2. Grafana
  3. Prometheus
  4. AWS CloudWatch
  5. ELK / OpenSearch
  6. Zabbix
  7. Nagios
  8. Datadog
  9. Other infrastructure/network monitoring platforms

Key Responsibilities During an Incident
  1. The NOC Engineer should be able to follow a structured process:
  2. Alert → Acknowledge → Validate → Troubleshoot → Identify Impact → Escalate → Monitor Resolution → Confirm Recovery → Document/RCA
  3. The candidate must be comfortable working under pressure during production outages and high-priority incidents.


Required Skills

CTI VoIP AWS Cloud SIP