Pass Perl Strings to C Safely with the perlguts API
You will leave a Perl scalar in the right form for a C function: a pointer plus an explicit byte length, with a deliberate choice between Perl's byte and UTF-8 views. This is the practical part of perlguts, the Perl API guide. Allow about twenty minutes if you already have an XS or embedding build; allow longer if you are still setting up the extension toolchain.
The route
Jump straight to the step you need, or tick off Done means at the end.
- 1. Confirm the API version you are reading
- 2. Treat an SV as a value with more than one representation
- 3. Choose the C view before extracting the buffer
- 4. Keep the pointer and length in separate statements
- 5. Respect the NUL-termination trap
- 6. Check definedness and truth separately
- 7. Verify the finished boundary
The examples target the installed Perl v5.38.2 and the matching perl-doc package. You need a C compiler and Perl development headers for a real build, but the checks in this guide only inspect the installed documentation and interpreter. Nothing here needs root, and no running service is changed.
1. Confirm the API version you are reading
Start by checking the interpreter and locating the local manual. This avoids quietly copying an API detail from a different Perl installation:
$ perl -v | sed -n '1,8p'
$ perldoc -l perlguts
/usr/share/perl/5.38/pod/perlguts.pod
The manual page installed with this system is generated for Perl v5.38.2. The broad shape of the API is old, but details can be version-specific. For example, perlguts says that retrieving a string from an integer no longer sets the scalar's POK flag from Perl 5.36.0 onwards. Do not use a different Perl's headers while testing code against this interpreter.
Checkpoint
The interpreter version and the documentation path should describe the same Perl installation. If perldoc cannot find the page, install or select the matching documentation package through your normal package-management process before relying on examples.
2. Treat an SV as a value with more than one representation
In the C API, an SV * is Perl's scalar value. It may represent an integer, unsigned integer, floating-point number, string, another scalar, or undef. The constructors make the intended initial value explicit:
SV *count = newSViv(42);
SV *ratio = newSVnv(0.5);
SV *label = newSVpvn(buffer, buffer_len);
Use a length-taking function when the data is already a buffer. newSVpvn accepts a STRLEN, so embedded NUL bytes are not mistaken for the end of the value. The shorter newSVpv form can calculate a length when its length argument is zero, but that calculation uses strlen. It therefore requires a NUL-terminated string with no earlier NUL byte.
Those are allocation examples, not complete ownership rules. An SV returned by newSV* normally needs the usual Perl reference-count or mortal-value handling in its surrounding API code. Do not paste an isolated constructor into an XS routine and assume the memory will manage itself.
3. Choose the C view before extracting the buffer
Perl can expose a scalar as bytes or as UTF-8. The choice depends on what the receiving C API means by a byte. For a binary protocol, digest input or file format, request bytes. For a C library that expects encoded text, request UTF-8 and make that encoding boundary explicit:
SV *sv = /* scalar received from Perl */;
STRLEN len;
char *ptr;
ptr = SvPVbyte(sv, len); /* byte string, length in bytes */
send_bytes(ptr, len);
ptr = SvPVutf8(sv, len); /* UTF-8 view, length in bytes */
send_text(ptr, len);
The macros write the resulting length into len; do not pass &len. If the length is irrelevant, the corresponding SvPVbyte_nolen or SvPVutf8_nolen form exists, but a C API that accepts a length is safer for arbitrary Perl strings.
Do not default to SvPV merely because the receiving parameter is char *. It exposes Perl's raw internal buffer, whose byte representation depends on the scalar's UTF-8 flag. The manual warns that using SvPV without checking that flag is almost certainly a bug when non-ASCII input is allowed. The explicit byte or UTF-8 variants state your intent and avoid making the next maintainer reverse-engineer it.
4. Keep the pointer and length in separate statements
Do not compress extraction into a call such as send_bytes(SvPVbyte(sv, len), len). The documented safe pattern assigns the pointer first, then passes both values:
SV *sv;
STRLEN len;
char *ptr;
ptr = SvPVbyte(sv, len);
send_bytes(ptr, len);
This is easier to review, and it avoids relying on evaluation order around a macro that can coerce or modify the scalar's representation. Keep len as the authoritative size. Do not replace it with strlen(ptr): Perl strings may contain NUL bytes.
Checkpoint
Inspect every call boundary. It should be obvious whether the callee receives bytes or UTF-8, and it should receive the length produced by the extraction macro. If the callee only accepts NUL-terminated text, verify that its contract really permits arbitrary Perl input. A length-aware adapter is safer than silently truncating at the first NUL.
5. Respect the NUL-termination trap
Perl commonly keeps a trailing NUL in string storage, but the API also permits arbitrary strings that contain NULs and may not be terminated in the way a particular C call expects. The trailing byte is not a substitute for a length: it is only safe when the receiving function's contract matches it.
When you grow an SV buffer yourself, account for the terminator explicitly. SvGROW increases allocated storage but does not automatically add space for the trailing NUL:
STRLEN old_len = SvCUR(sv);
char *dst = SvGROW(sv, old_len + wanted + 1);
STRLEN added = read(fd, dst + old_len, wanted);
dst[old_len + added] = '\0';
SvCUR_set(sv, old_len + added);
SvUTF8_off(sv);
SvSETMAGIC(sv);
This pattern is only suitable when the bytes you read are meant to be a byte string. It also assumes that errors and short reads are handled in the real code. If you already have the complete data in memory, prefer sv_setpvn or sv_catpvn; the simpler length-taking functions are less error-prone than manual buffer bookkeeping.
6. Check definedness and truth separately
Undef is a Perl value, not a null pointer. Use SvOK(sv) to test whether a scalar is defined and SvTRUE(sv) to test its truth value:
if (!SvOK(sv)) {
/* Perl passed undef */
} else if (SvTRUE(sv)) {
/* a defined, true value */
}
Do not compare an arbitrary SV pointer with &PL_sv_undef. A scalar variable assigned undef is not necessarily that singleton SV. Likewise, (SV *)0 is a null pointer, not Perl's undef value. If your API needs to return undef, use the Perl undef value supplied by the API, and keep null pointers for actual failure or absence.
7. Verify the finished boundary
Before compiling or shipping the extension, search the C code for the places where Perl values cross into another library:
$ rg -n 'SvPV|SvGROW|strlen|SvOK|SvTRUE' path/to/extension
For each match, record the intended encoding, the byte length source and the behaviour for undef. Then build and run the extension's existing tests with the same perl reported in step 1. A successful build does not prove that a non-ASCII string or embedded NUL is handled correctly, so include both cases in tests when the receiving library permits them.
Done means
- You confirmed the Perl interpreter and
perlgutsdocumentation versions match. - Each extracted string uses an explicit byte or UTF-8 API and passes its returned length.
- Your code does not use
strlenfor arbitrary Perl strings. - Manual buffer growth reserves and writes the trailing NUL when the surrounding API requires it.
- Undef is checked with
SvOK, not confused with a null pointer or a particular SV address.