Route Packets into a VXLAN with tc tunnel_key
You will build a traffic-control filter that matches incoming ICMP packets, attaches VXLAN tunnel metadata, and redirects the packets to a shared VXLAN device for encapsulation. The same pattern can carry Geneve or ERSPAN metadata, but this guide keeps the working example to VXLAN. Allow about 10 minutes if the tunnel device already exists.
The route
Jump straight to the step you need, or tick off Done means at the end.
Before you start
This is an administrative networking change. The commands that create or remove qdiscs and filters need root privileges, so the examples use sudo. Run them on a test interface or during a maintenance window: a filter can change which packets leave a host.
You need:
- the
tccommand from iproute2; - an ingress interface, called
eth0below; - a shared IP tunnel device, called
vxlan0below; and - an outer source and destination address that are valid for your network.
On the machine used for this guide, iproute2 is version 6.1.0-1ubuntu6.4 and tc reports iproute2-6.1.0. The exact accepted options can vary with iproute2, so check the installed manual page with man 8 tc-tunnel_key before moving this to another distribution.
Checkpoint: understand the hand-off
tunnel_key does not create a VXLAN device and does not send a packet by itself. In set mode it places outer source and destination addresses and a tunnel ID in packet metadata. A following mirred redirect sends that packet to the shared tunnel device, which uses the metadata when it encapsulates the packet.
For the example, the tunnel ID is 11, which is a typical VXLAN network identifier value. The source and destination below are placeholders: replace them with addresses belonging to the two tunnel endpoints.
1. Add an ingress hook
Ingress filters need an ingress qdisc. This command changes the interface configuration and requires elevation:
sudo tc qdisc add dev eth0 handle ffff: ingress
Verify that the hook exists:
sudo tc qdisc show dev eth0
You should see an ingress qdisc with handle ffff:. If one already exists, do not run the add command again; continue to the filter step. An "already exists" error is normally a sign that the hook is already present, not that the interface is unusable.
2. Set metadata and redirect matching packets
Add a flower filter for ICMP over IPv4. The first action sets the metadata. The second action redirects the packet to vxlan0 for encapsulation:
sudo tc filter add dev eth0 protocol ip parent ffff: \
flower ip_proto icmp \
action tunnel_key set \
src_ip 192.0.2.10 \
dst_ip 198.51.100.20 \
id 11 \
action mirred egress redirect dev vxlan0
The addresses use documentation ranges. They are safe as examples, but they are not a working tunnel endpoint pair. Substitute real addresses and make sure the tunnel device is configured to use the metadata. id is the tunnel key, commonly the VXLAN VNI. The source and destination are the outer header addresses, not the original packet's addresses.
Inspect the installed filter:
sudo tc filter show dev eth0 ingress
Look for a flower filter followed by tunnel_key and mirred. Packet counters are useful evidence that the match is being hit, but they only become meaningful after matching traffic arrives. If the counters stay at zero, check the protocol, interface direction, and whether the packet is actually ICMP.
3. Choose checksum behaviour deliberately
The default is csum: the outer UDP checksum is calculated and included. Keep that default unless the tunnel design specifically requires zero checksums. You can state it explicitly in the action:
action tunnel_key set src_ip 192.0.2.10 dst_ip 198.51.100.20 id 11 csum
nocsum makes the outer UDP checksum zero. With IPv6, the other endpoint must be configured to accept that, and a zero checksum gives weaker protection against corrupted packets. Do not use nocsum as a troubleshooting shortcut. Change the filter deliberately, then verify the resulting packets with a capture on the appropriate interface.
4. Remove the test configuration
Removing the filter is the first rollback action. The filter was added without an explicit handle, so delete the matching filter by repeating its selector:
sudo tc filter del dev eth0 protocol ip parent ffff: \
flower ip_proto icmp
Confirm that it is gone:
sudo tc filter show dev eth0 ingress
If this ingress qdisc was created only for this test, remove it after removing the filter:
sudo tc qdisc del dev eth0 ingress
Do not delete the qdisc if other ingress filters depend on it. Deleting the qdisc removes its filters and can disrupt unrelated traffic processing.
When to use unset
Metadata is released automatically when packet processing finishes, so most one-tunnel paths need only set. Use unset when a packet has been decapsulated by one tunnel and is about to be redirected towards another tunnel. It clears the metadata from the first tunnel so that stale outer addresses or a stale key are not reused.
The action belongs before the redirect in that path:
action tunnel_key unset \
action mirred egress redirect dev tunl1
In a real rule, add the appropriate flower encapsulation matches, such as the source, destination and tunnel key received from the first tunnel. Treat the example as the action sequence, not as a complete production selector.
Common traps
- Expecting a tunnel to appear.
tunnel_keysupplies metadata; it does not createvxlan0. Check the tunnel device separately. - Using only one action. Without the
mirredredirect to a shared tunnel device, the metadata has no tunnel device to consume it. - Reusing an existing ingress hook. An existing
ffff:qdisc may contain other filters. Inspect it before changing or deleting anything. - Confusing the tunnel ID with an IP address.
idis the tunnel key, such as a VXLAN VNI;src_ipanddst_ipdescribe the outer IP header. - Testing the wrong traffic. The example matches ICMP only. Generate a controlled ping or broaden the flower match only after deciding which packets should be encapsulated.
Done means
tc qdisc show dev eth0shows the intended ingress hook.tc filter show dev eth0 ingressshowstunnel_key setfollowed by the redirect.- The tunnel device exists and is configured for the chosen endpoints and key.
- Filter counters increase only for the traffic you intended to match.
- You have recorded the exact delete commands, or removed the test filter and qdisc safely.