Founder & CEO of Unvariance: empowering SREs to deliver exceptional performance and efficiency in the cloud.
Original founder and maintainer of OpenTelemetry Network (opentelemetry-network)
Previously: Founder and CEO of Flowmill (acquired by Splunk)
PhD in Networks and Distributed Systems, MIT
I currently lead Unvariance, where we’re working to solve one of cloud computing’s fundamental challenges: performance variability caused by resource contention between co-located workloads. Our initial focus is on memory noisy neighbor detection and mitigation, building on experience with low-overhead monitoring and real-time system analysis.
At Flowmill, we built tools to help SREs handle production incidents. The technology enabled faster incident resolution by monitoring service dependencies, helping teams identify problem sources within seconds and reduce war room staffing. After Splunk’s acquisition, we open-sourced the technology to the Cloud Native Computing Foundation, where it continues to evolve as the OpenTelemetry Network project, achieving quarter-percent CPU utilization and 0.5% network overhead while providing rich contextual data about service interactions with no sampling and no code changes.
I received my Ph.D at MIT CSAIL‘s Networks and Mobile Systems group, advised by Hari Balakrishnan and Devavrat Shah, with the thesis “Centralized performance control for datacenter networks“, during which we collaborated with Microsoft Research (2011) and Facebook (2013-2017). I had previously spent 7 years in communication systems R&D and HPC algorithm development as an officer in an army technological unit.
The PhD research revolved around enabling fast detection of and reaction to undesirable incidents in datacenter and cloud networks, by designing extremely fine granulrity, low overhead, low latency monitoring, processing, and control of service interactions. The systems produced mostly controlled network transfers, as this use-case provides extreme challenges for the technology. Fastpass aims for high utilization with zero queueing: a logically centralized arbiter controls and orchestrates all network transfers. Flowtune assigns shares of network throughput to pairs of applications according to organizational policy, maximizing the organization’s utility.
Other research deals with rateless error correcting codes for wireless networks: Spinal Codes (w/source code) are efficient, high-performance error correction codes, especially suited for analog channels.
Selected Publications
- J. Perry, H. Balakrishnan, and D. Shah. Flowtune: Flowlet Control for Datacenter Networks, NSDI 17.
- J. Perry, A. Ousterhout, H. Balakrishnan, D, Shah, H. Fugal, Fastpass: A Centralized “Zero-Queue” Datacenter Network, SIGCOMM 2014.
- J. Perry, P. Iannucci, K. Fleming, H. Balakrishnan, D, Shah, Spinal Codes, SIGCOMM 2012.
- J. Perry, H. Balakrishnan, D. Shah, Rateless Spinal Codes, HotNets 2011.
More publications
- P. Iannucci, J. Perry, H. Balakrishnan, D. Shah, No Symbol Left Behind: A Link-Layer Protocol for Rateless Codes, MobiCom 2012.
- D. Shah, J. Perry, P. Iannucci, H. Balakrishnan, De-randomizing Shannon: The Design and Analysis of a Capacity-Achieving Rateless Code, Manuscript in preparation/submission.
- P. Iannucci, K. Fleming, J. Perry, H. Balakrishnan, D. Shah, A Hardware Spinal Decoder, ANCS 2012.
Teaching
Spring 2014: 6.824 Distributed Systems
Spring 2013: 6.829 Computer Networks
