# Python Strings: F-Strings, Methods and the strip() Trap

> Python strings in practice: f-strings, slicing, the methods worth memorising, why += loops slow down, and the encoding default that breaks scripts on Windows.

- **Author:** Ayushi Kulshreshta — AI Engineer (https://www.1stepgrow.com/authors/ayushi-kulshreshta/)
- **Published:** Jul 24, 2026 · **Updated:** Sep 17, 2026
- **Topic:** Python & Programming · **Format:** Guide · **Read time:** 9 min
- **Canonical URL:** https://www.1stepgrow.com/articles/python-strings/

## Key takeaways

- F-strings are the default in modern Python — clearer and faster than format() or % formatting.
- Strings are immutable, so building one in a loop with += can degrade to quadratic time. Use join().
- str.strip() removes any of the given characters, not a suffix — use removeprefix/removesuffix for that.
- Always specify encoding='utf-8' when opening text files; the default depends on the platform before Python 3.15.

Python strings are immutable sequences of Unicode characters: you can read, slice and search them, but every method hands back a new string. That one rule explains most string bugs, from `+=` loops that slow to a crawl to `"telecom.com".strip(".com")` returning `'tele'`.

If you clean CSV columns, parse log lines or format reports, this guide covers the parts that cause real friction: f-strings and format specs, slicing, the methods worth memorising, and the encoding default that lets a script pass on Linux and fail on Windows. It follows [Python numeric data types](https://www.1stepgrow.com/articles/python-numeric-data-types) in the Python series. The full method list lives in the official [string methods reference](https://docs.python.org/3/builtins/stdtypes.html#string-methods).

## How do f-strings work?

An f-string is a string literal prefixed with `f`, and anything inside braces is evaluated as an expression. It is the modern default, and the only formatting you need for new code:

```python
name = "Ananya"
score = 0.8734
count = 1_234_567

f"Hello, {name}"                    # 'Hello, Ananya'
f"Score: {score:.1%}"               # 'Score: 87.3%'
f"Count: {count:,}"                 # 'Count: 1,234,567'
f"Pi: {3.14159:.2f}"                # 'Pi: 3.14'
f"Hex: {255:#x}"                    # 'Hex: 0xff'
f"Padded: {42:>8}"                  # 'Padded:       42'
f"Centred: {'hi':^10}|"             # 'Centred:     hi    |'
```

The codes after the colon come from the [format specification mini-language](https://docs.python.org/3/library/string.html#formatspec), which is worth bookmarking. Expressions also work inside the braces:

```python
items = [1, 2, 3]
f"There are {len(items)} items, total {sum(items)}"   # 'There are 3 items, total 6'
```

Since Python 3.12, f-strings can reuse the enclosing quote character inside the braces, as described in [PEP 701's What's New entry](https://docs.python.org/3/whatsnew/3.12.html#pep-701-syntactic-formalization-of-f-strings). On older versions, this line is a syntax error:

```python
cities = ["Mumbai", "Delhi"]
f"Cities: {", ".join(cities)}"      # 'Cities: Mumbai, Delhi' (3.12+)
```

The debugging form is genuinely useful and underused — a trailing `=` prints the expression alongside its value:

```python
x = 42
f"{x=}"                # 'x=42'
f"{score * 100=:.1f}"  # 'score * 100=87.3'
```

Multi-line output with alignment:

```python
rows = [("Mumbai", 20.4), ("Delhi", 32.9), ("Pune", 7.4)]
for city, pop in rows:
    print(f"{city:<10} {pop:>6.1f}M")
```

```text
Mumbai       20.4M
Delhi        32.9M
Pune          7.4M
```

Python 3.14 also added template strings (t-strings), which use a `t` prefix and return a `Template` object instead of a `str`. They exist for libraries that need to process the parts safely, such as HTML or SQL escaping; the [Python 3.14 release notes](https://docs.python.org/3/whatsnew/3.14.html#pep-750-template-string-literals) cover them. For everyday formatting, however, f-strings remain the tool.

## Which string formatting style should you use?

| Style | Example | Use it for |
|---|---|---|
| f-string | `f"{name}: {score:.1f}"` | almost everything |
| `str.format()` | `"{}: {:.1f}".format(name, score)` | templates stored apart from their data |
| `%` formatting | `"%s: %.1f" % (name, score)` | reading legacy code and `logging` calls |

## Indexing and slicing Python strings

Strings are sequences, so indexing and slicing work exactly as they do for lists:

```python
s = "Python"
s[0]        # 'P'
s[-1]       # 'n'
s[1:4]      # 'yth'
s[::-1]     # 'nohtyP' - reversed
s[10:]      # '' - slices never raise IndexError
```

However, you cannot assign to an index, because strings are immutable:

```python
s[0] = "J"
# TypeError: 'str' object does not support item assignment
"J" + s[1:]  # 'Jython' - build a new string instead
```

## Why is building a string with += slow?

Because strings are immutable, every "modification" creates a new string and copies the old contents into it:

```python
s = "hello"
s.upper()      # returns 'HELLO'
print(s)       # hello - the original is unchanged
```

Which makes this pattern potentially quadratic:

```python
words = ["data", "science", "course"]

# bad - O(n²) in general
result = ""
for word in words:
    result += word + " "

# good - O(n)
result = " ".join(words)     # 'data science course'
```

With a thousand words the difference is invisible. With a million, it can be the difference between instant and a coffee break. CPython sometimes resizes the string in place for `+=`, so you may not see the slowdown locally, but that is an implementation detail you should not rely on. Therefore `join` remains the idiomatic answer whenever you are assembling a string from parts.

For building with logic, collect then join:

```python
records = [{"name": "a", "value": 1, "valid": True},
           {"name": "b", "value": 2, "valid": False}]
parts = []
for row in records:
    if row["valid"]:
        parts.append(f"{row['name']}: {row['value']}")
output = "\n".join(parts)    # 'a: 1'
```

## Which string methods are worth memorising?

A small set of methods covers most everyday text work: trimming, case changes, splitting and joining, searching, and testing what a string contains.

```python
s = "  Data Science Course  "

s.strip()              # 'Data Science Course'
s.lower()              # '  data science course  '
s.title()              # '  Data Science Course  '
s.replace("Course", "Program")   # '  Data Science Program  '

"a,b,c".split(",")     # ['a', 'b', 'c']
"a b  c".split()       # ['a', 'b', 'c'] - splits on any whitespace run
"a,b,c".split(",", 1)  # ['a', 'b,c'] - limit the splits

"hello".startswith("he")     # True
"file.csv".endswith(".csv")  # True
"abc" in "xxabcxx"           # True

"hello".find("l")      # 2, or -1 if absent
"hello".index("l")     # 2, raises ValueError if absent
"hello".count("l")     # 2
```

Case-checking and padding:

```python
"abc123".isalnum()     # True
"123".isdigit()        # True
"abc".isalpha()        # True
"  ".isspace()         # True

"5".zfill(3)           # '005'
"x".ljust(5, ".")      # 'x....'
```

For quick reference, these are the methods you will reach for most often:

| Method | What it does | Returns |
|---|---|---|
| `strip()` / `lstrip()` / `rstrip()` | removes a *set* of characters from the ends | new `str` |
| `removeprefix()` / `removesuffix()` | removes one exact substring, once | new `str` |
| `split()` / `join()` | breaks apart / glues together | `list` / `str` |
| `replace(old, new)` | substitutes every occurrence | new `str` |
| `find()` / `index()` | locates a substring | `int` (`-1` or `ValueError` if absent) |
| `startswith()` / `endswith()` | checks the ends; also accepts a tuple | `bool` |
| `casefold()` | aggressive lowercase for comparisons | new `str` |

## Why does strip() remove more than you expect?

`strip` takes a **set of characters**, not a substring. The [`str.strip` documentation](https://docs.python.org/3/builtins/stdtypes.html#str.strip) says so directly, yet the name still misleads people:

```python
"telecom.com".strip(".com")     # 'tele'   <- not what you wanted
```

It removed every leading and trailing character in `{'.', 'c', 'o', 'm'}`, so it kept eating past `.com` into `telecom`. With a name like `example.com` the bug stays hidden, because `e` is not in the set — which is exactly why it survives testing.

Use the explicit methods instead:

```python
"telecom.com".removesuffix(".com")     # 'telecom'
"www.example.com".removeprefix("www.") # 'example.com'
```

These were added in Python 3.9 and are the correct tool. Before that, the idiom was a conditional slice, which is why the bug was so common.

## How do you avoid encoding errors with Python strings?

Encoding is the source of most real-world string pain in data work, and the fix is to be explicit.

```python
# always be explicit
with open("data.csv", encoding="utf-8") as f:
    text = f.read()

# for messy sources that may contain invalid bytes
with open("data.csv", encoding="utf-8", errors="replace") as f:
    text = f.read()
```

Without an explicit encoding, [`open()`](https://docs.python.org/3/builtins/functions.html#open) uses a platform-dependent locale encoding. For instance, code that works on your Linux CI can fail on a colleague's Windows laptop with a `UnicodeDecodeError` on a single accented character. [PEP 686](https://peps.python.org/pep-0686/) makes UTF-8 mode the default from Python 3.15. As of September 2026, however, the stable documentation still covers Python 3.14, so until every machine you support runs 3.15 or later, keep passing `encoding="utf-8"`.

Encoding and decoding manually:

```python
text = "café"
data = text.encode("utf-8")     # b'caf\xc3\xa9' - bytes
data.decode("utf-8")            # 'café' - back to str
```

The mental model: `str` is text, while `bytes` is what goes on disk or over the network. Encode to serialise, decode to read.

Watch out for characters where visual length differs from `len`:

```python
len("café")        # 4
len("👨‍👩‍👧")      # 5 - one visible glyph, several code points
```

If you are truncating user-facing text, this matters.

## How do you clean messy text columns?

Chain a few methods in a small function: normalise odd whitespace, trim, then lowercase. This pattern shows up constantly in data work:

```python
def clean(value: str) -> str:
    return (
        value.replace(" ", " ")     # non-breaking space to normal space
             .strip()
             .lower()
    )

raw_columns = ["  Customer ID ", "Order Date"]
columns = [clean(c) for c in raw_columns]   # ['customer id', 'order date']
```

Non-breaking spaces from spreadsheet exports are a genuine and infuriating cause of failed joins, because the values look identical and are not. Note the order, too: replace them before `strip()`, since the replacement can create new leading or trailing spaces.

## Common mistakes with Python strings

- **Using `strip()` to remove a suffix.** Use `removesuffix()` instead.
- **Forgetting that methods return new strings.** `s.upper()` on its own line does nothing useful; assign the result.
- **Building long strings with `+=` in a loop.** Collect parts and `join` them.
- **Opening files without `encoding=`.** The default differs between machines.
- **Comparing user input case-sensitively.** Use `casefold()` on both sides.

For anything more complex than fixed substitutions, use [regular expressions](https://www.1stepgrow.com/articles/regular-expressions-in-python). Next in the series, [Python operators](https://www.1stepgrow.com/articles/python-operators) covers `in`, `+` and `*` on strings alongside everything else.



## Related reading

[Regular expressions Part 1](https://www.1stepgrow.com/articles/regular-expressions-in-python) covers pattern matching. [Python basics](https://www.1stepgrow.com/articles/python-basics) covers the core types.

## Frequently asked questions

### Why is += slow for building strings?

Strings are immutable, so each += has to create a new string and copy both operands into it. Repeating that n times is O(n²) in the general case. CPython sometimes optimises the pattern in place, but you cannot rely on it across interpreters. Collect the pieces in a list and call ''.join(pieces) instead, which is O(n).

### What is the difference between strip and removeprefix?

strip('abc') removes any leading or trailing characters that appear in the set {a, b, c}, repeatedly, until it meets a character outside the set. removeprefix('abc') removes exactly that one substring, once, if present. Using strip to remove a file extension or domain suffix is a classic bug, and removeprefix and removesuffix have existed since Python 3.9.

### Which string formatting should I use?

Use f-strings for essentially everything in Python 3.6 and later. They are readable, fast, and support format specs and arbitrary expressions, and since Python 3.12 they can even reuse the same quote character inside the braces. Keep str.format() for templates defined separately from their data, and avoid % formatting in new code.

### How do I avoid encoding errors when reading files?

Pass encoding='utf-8' explicitly to open(). Before Python 3.15 the default comes from the platform locale, so code that works on Linux can fail on Windows. For messy sources, errors='replace' substitutes a replacement character for unreadable bytes instead of raising UnicodeDecodeError, which lets you inspect the damage rather than crash.

---
_Source: 1stepGrow (https://www.1stepgrow.com/articles/python-strings/). Cite with the title, "1stepGrow" and a link._
