Strings & Runes
Understand that Go strings are UTF-8 byte sequences, what the rune type represents, how to correctly iterate Unicode text, and the strings package basics.
Introduction
Text handling is one of the areas where Go's design choices are both distinctive and occasionally surprising to newcomers. A Go string is not an array of characters — it's an immutable sequence of bytes, conventionally interpreted as UTF-8 encoded text. Understanding this distinction, along with the rune type, is essential for handling Unicode text correctly.
- Why Go strings are byte sequences, not character arrays.
- What the rune type is and when to use it.
- How range correctly iterates Unicode code points in a string.
- The difference between byte-based and rune-based indexing.
- Common functions from the strings package.
Strings Are Byte Sequences
A Go string is fundamentally a read-only slice of bytes. When the string contains only ASCII characters, each byte corresponds neatly to one character. But for non-ASCII text — accented letters, emoji, Chinese characters — a single visible character can be encoded as multiple bytes in UTF-8.
package main
import "fmt"
func main() { greeting := "Café"
fmt.Println("string:", greeting) fmt.Println("len (bytes):", len(greeting))
for i := 0; i < len(greeting); i++ { fmt.Printf("byte %d: %d (%q)\n", i, greeting[i], string(greeting[i])) }}Click Run to see what this code prints.
Notice len("Café") is 5, not 4 — the "é" character takes two bytes in UTF-8. Indexing a string with s[i] gives you a single byte, which is why the last two bytes print as garbled characters when converted individually.
The Rune Type
A rune is Go's type for a single Unicode code point — it's just an alias for int32. Where a byte represents one raw byte of UTF-8 encoding, a rune represents one full character, however many bytes it took to encode. Converting a string to []rune gives you a slice where each element is exactly one character.
package main
import "fmt"
func main() { greeting := "Café"
runes := []rune(greeting) fmt.Println("rune count:", len(runes))
for i, r := range runes { fmt.Printf("rune %d: %c (code point %d)\n", i, r, r) }}Click Run to see what this code prints.
Iterating a String With Range
When you range directly over a string (without converting it to []rune first), Go automatically decodes UTF-8 for you: each iteration gives you the byte index where a character starts, and the decoded rune itself — not a raw byte. This is almost always what you want when processing text character by character.
package main
import "fmt"
func main() { word := "Café"
for index, char := range word { fmt.Printf("byte index %d: %c\n", index, char) }}Click Run to see what this code prints.
Notice the index jumps from 3 straight to the end — that's because "é" occupies byte positions 3 and 4, and range correctly reports the starting byte of each full character rather than stepping one raw byte at a time.
Byte Indexing vs Rune Indexing
Because s[i] indexes bytes, not characters, you cannot safely use it to grab the "nth character" of a string containing multi-byte characters. If you need positional character access, convert to []rune first and index into that instead.
package main
import "fmt"
func nthChar(s string, n int) rune { runes := []rune(s) if n < 0 || n >= len(runes) { return 0 } return runes[n]}
func main() { word := "Café" fmt.Printf("4th character: %c\n", nthChar(word, 3))}Click Run to see what this code prints.
The strings Package
The standard library's strings package provides a large set of ready-made functions for common text operations — searching, splitting, joining, replacing, trimming, and case conversion.
package main
import ( "fmt" "strings")
func main() { sentence := " Go is fun, Go is fast "
fmt.Println(strings.TrimSpace(sentence)) fmt.Println(strings.ToUpper("hello")) fmt.Println(strings.Contains(sentence, "fast")) fmt.Println(strings.Count(sentence, "Go")) fmt.Println(strings.Replace(sentence, "Go", "Golang", -1))
parts := strings.Split("a,b,c", ",") fmt.Println(parts)
joined := strings.Join([]string{"x", "y", "z"}, "-") fmt.Println(joined)}Click Run to see what this code prints.
Common Mistakes
- Using len(s) to mean "number of characters" — it counts bytes, which differs from character count for non-ASCII text.
- Indexing a string with s[i] expecting a character — it returns a single byte, not a rune.
- Manually looping with for i := 0; i < len(s); i++ over text that might contain multi-byte characters, instead of using range.
- Trying to mutate a string in place — Go strings are immutable; you must build a new string instead.
- Forgetting to import "strings" (the package) when reaching for functions like Split or Join.
Best Practices
- Use range over a string when you need to process it character by character — it decodes UTF-8 correctly for you.
- Convert to []rune only when you need random-access indexing or a true character count.
- Reach for the strings package before writing manual string manipulation loops — it's well-tested and often faster.
- Use a strings.Builder for efficiently constructing a string incrementally in a loop, rather than repeated concatenation with +.
- Be mindful of UTF-8 when working with any user-supplied or internationalized text.
Frequently Asked Questions
byte is an alias for uint8, representing one raw byte of data. rune is an alias for int32, representing one full Unicode code point, which might be encoded as 1 to 4 bytes in UTF-8.
Use len([]rune(s)) or utf8.RuneCountInString(s) from the unicode/utf8 package — len(s) alone gives you the byte count instead.
No, strings are immutable in Go — that line won't even compile. Convert to []byte or []rune, modify that, then convert back to a new string.
Not guaranteed by the type system — a string can technically hold arbitrary bytes. By convention and in standard library functions, though, strings are treated as UTF-8.
Key Takeaways
- Go strings are immutable sequences of bytes, conventionally UTF-8 encoded text.
- A rune represents one full Unicode code point and is an alias for int32.
- range over a string decodes UTF-8 automatically, giving you the byte index and the rune.
- Indexing a string with s[i] returns a raw byte, not a character — convert to []rune for character-level access.
- The strings package provides efficient, well-tested functions for nearly every common text operation.
Summary
Understanding the byte-versus-rune distinction is one of the most important mental models for writing correct, internationalization-safe Go code. Get comfortable with range over strings and the strings package, and you'll rarely run into text-handling bugs. Next, you'll shift from data types to behavior, starting with how to declare and use functions.