What an internet exchange does
Every computer on the internet, whether it is your phone or one of the many servers used by Google, has to know how to reach every other one. Your phone must know how to get to Google, and Google must know how to send the reply back. Because there are a great many networks in the world, it would be impossible for all of them to connect directly to each other. So networks run a protocol called BGP, which each network uses to tell its neighbours which destinations it can reach. Those neighbours pass the message on, until the whole world knows the route.
Every extra network a packet has to cross costs time, so direct links are usually better. Networks of a similar size often exchange traffic for free because both sides benefit. But direct links do not scale: with 183 networks here, a newcomer wanting to reach all of them would need 183 separate connections.
An internet exchange solves that. Everyone makes one connection to a shared switch, and from there can reach everyone else on it. That is all an IX fundamentally is: a layer-2 fabric plus the agreements around it.
The catch is that a traditional exchange is generally only available to networks with equipment in the same building, and the building is expensive. That is perfectly reasonable for a business, and prohibitive for a hobbyist. EVIX exists to remove that barrier: the fabric is built from tunnels, so anyone with an ASN can join from anywhere, at no cost, and learn how a real exchange behaves. The skills transfer directly to a job. It has been just as good a learning exercise for those of us running it, because operating something that 183 people depend on raises problems that never come up in a lab.
Our goals
- Provide the lowest possible barrier to entry for hobbyists and learners who want experience with something as close to a real internet exchange as we can make it.
- Automate as much of the operation as we can, both to keep the workload on volunteers sustainable and because the automation is itself the interesting part.
- Be a useful reference for other exchanges. Our tooling is public, and we are happy for people to learn from it.
Why virtual, and the objections to it
The obvious question is why we did not build a free exchange in a datacentre, the way FCIX did. The answer is cost. When Bryce started EVIX there was no way to afford a full rack, even in the very cheap facility FCIX is in, and FCIX requires a whole rack to join; in most other facilities a cross-connect carries a monthly fee on top. Starting that way would have meant sponsorship from day one, and continuing would have meant charging real money. The only way to serve a large number of people who cannot afford a rack, and in many cases cannot afford a VPS either, was to make the fabric virtual.
Not everyone agrees that virtual exchanges should exist, and we would not claim every one of them is justified. The usual suggestion is that we should use DN42 or build a lab instead. DN42 is genuinely good and we recommend it to anyone without an ASN yet. But some things cannot be learned from either.
Running this means working as a team, which is much harder than working alone. It means supporting platforms we would never have chosen ourselves, because members turn up with them. It means automating things nobody planned for, because the member count made the manual version impossible. It means living with the technical debt of a system that has been running for years, and with the fact that a few hundred people depend on it working. That is much closer to a job than a hobby project is, and that experience alone justified building it.
We are happy to discuss this properly, so do get in touch. There is usually a compromise available once both sides are actually talking.
Supported connection types
Because every member sets their own routing policy towards every other member, the fabric has to carry layer-2 frames, not just IP packets. That is the layer Ethernet and MAC addresses live at. Plain Generic Routing Encapsulation (GRE) is therefore not supported, since it can only carry layer-3 packets. Its layer-2 variant GRETAP is supported.
| Type | Good for | Notes |
|---|---|---|
| VXLAN | Most Linux and BSD routers | Our default recommendation. UDP-based, so it survives NAT. |
| GRETAP | Routers without VXLAN support | Needs a clean public IPv4 endpoint. GRE carries no port numbers and cannot traverse NAT, so use VXLAN if there is NAT in the path. |
| EoIP | MikroTik RouterOS | Easiest option on RouterOS. |
| OpenVPN | VyOS, and hosts behind awkward NAT | Our server requires LZO compression. See the example configuration. |
| ZeroTier | Getting started quickly | Least predictable performance of the tunnel types. |
| Local Ethernet | A VM or port in a facility we are in | No tunnel at all. This is the right choice for a VM from one of our hosting partners. |
| Custom layer-2 transit | A VLAN handed to us by a carrier | Possible at Amsterdam, Zurich and Frankfurt. Ask us. |
We are open to other layer-2 protocols. Mention what you have in mind in your ticket.
How it is built
- The switch fabric is a mesh of VXLAN tunnels between the sites, bridged into one broadcast domain per site. Peer tunnels terminate on the nearest site and join that bridge.
- Two route servers run BIRD 2, with configuration generated by ARouteServer. RPKI data reaches them from a local StayRTR instance over RTR.
- Peer records, address allocations and monitoring state live in a MySQL database. Provisioning, tunnel creation and route-server configuration are generated from it by Bash, jq and Python, and distributed with Ansible.
- All of our servers run Ubuntu Server.
- Monitoring is LibreNMS and SmokePing; the looking glass at lg.evix.org is generated from the live route-server state. Support runs through an osTicket helpdesk.
- This website is hand-written HTML and SCSS with a small Jinja build step, served by OpenLiteSpeed. The peers list and the headline figures on the homepage are generated from the database at build time.
- Configuration lives in Git on a GitLab server run by one of our admins. An older snapshot of it is public on GitHub; see using our code for what that is and is not.
Using our code
A good deal of our tooling is public on GitHub. Be aware of what it is: a snapshot that stopped updating in April 2022, when the mirroring broke and nobody noticed. It is still a useful read, and the ideas in it are still the ideas we run on, but it does not match what is deployed today. It was never a turnkey solution either, and parts of it are under-documented, so treat it as a reference for people who have hit similar problems rather than something to deploy. The jq scripts for parsing BIRD output are a good example of the useful bits.
Because that repository is dormant, issues and pull requests opened there are not seen. If you want to tell us something about the code, ask how a piece of it works, or contribute, please email us instead and we will answer.
We are glad to see our snippets reused. We do ask that you not copy the whole thing wholesale, in particular this website's design and text, or our complete set of scripts as a unit. If you are unsure, email us and we will happily talk it through, and we can usually explain the parts that are not documented well.
Donating
EVIX has no fees and no budget. Everything runs on donated infrastructure, so what helps most is donated resources rather than money: hardware, rack space, transit, or a VM for a new location in a region where we are thin. We are currently interested in a presence in Asia, Africa, South America, or the east coast of North America.
If you want to contribute money or hardware, or you have an idea for a location, email us and we will sort it out with you.
Contact
To join, use the peering request form. For anything else, email helpdesk@ev-ix.org, which opens a ticket, or find us in #evix-support on Discord.