Swift Markdown Kit · DOCS

Streaming & Chat

Streaming & Chat

This is the page to read if you are building a chat UI on top of a language model. Everything here is the default behaviour unless stated otherwise.

The whole integration

One MarkdownChatRenderView renders the entire conversation — one WKWebView for the whole screen, not one per message. Hand it your messages and let it diff:

import SwiftUI
import SwiftMarkdownKit

@MainActor
@Observable
final class ConversationModel {
    var messages: [MarkdownChatMessage] = []

    func send(_ prompt: String, using client: MyStreamingClient) async {
        messages.append(MarkdownChatMessage(role: .user, markdown: prompt))

        let replyID = UUID().uuidString
        messages.append(
            MarkdownChatMessage(id: replyID, role: .assistant, markdown: "", status: .streaming)
        )

        for await token in client.stream(prompt) {
            guard let index = messages.firstIndex(where: { $0.id == replyID }) else { return }
            messages[index].markdown += token
        }

        if let index = messages.firstIndex(where: { $0.id == replyID }) {
            messages[index].status = .completed
        }
    }
}

struct ConversationScreen: View {
    @State private var model = ConversationModel()

    var body: some View {
        MarkdownChatRenderView(messages: model.messages)
    }
}

Appending to markdown is all it takes. The view notices that a message grew rather than changed, and streams only the delta into the existing bubble — nothing is re-parsed or repainted from scratch, so the total cost of rendering a whole answer is linear in its length.

Set status to .completed when the stream ends. That is what stops the typing indicator and closes the reasoning fold.

Driving it from UIKit

For UIKit — or for any case where you would rather push than diff — talk to the view controller directly:

import UIKit
import SwiftMarkdownKit

final class ConversationViewController: UIViewController {
    private let chat = MarkdownChatRenderViewController()

    override func viewDidLoad() {
        super.viewDidLoad()

        addChild(chat)
        chat.view.frame = view.bounds
        chat.view.autoresizingMask = [.flexibleWidth, .flexibleHeight]
        view.addSubview(chat.view)
        chat.didMove(toParent: self)
    }

    func begin(replyTo prompt: String) {
        chat.appendMessage(MarkdownChatMessage(role: .user, markdown: prompt))
        chat.appendMessage(
            MarkdownChatMessage(id: "reply", role: .assistant, markdown: "", status: .streaming)
        )
    }

    func receive(_ token: String) {
        chat.appendChunk(token, to: "reply")
    }

    func finish() {
        chat.finishMessage("reply")
    }
}

appendChunk(_:to:isFinal:) is the hot path and is safe to call for every token. Tokens are gathered for roughly one frame before crossing into the web view, so a model producing forty tokens a second costs a couple of crossings, not forty. Text still appears at the same rate — see chunkCoalescingInterval if you want to turn that off.

Reasoning

Reasoning models publish their chain of thought in one of two shapes. The SDK handles both, and you do not have to know which one you are getting.

Inline — the thinking arrives inside the content stream wrapped in <think> … </think>, which is what DeepSeek-R1, Qwen, GLM and most locally served models do. Pass it through with everything else:

chat.appendChunk("<think>weighing the options</think>", to: "reply")
chat.appendChunk("The answer is 42.", to: "reply")

A marker split across two chunks — <thi then nk> — is still recognised.

A separate channel — the provider streams reasoning in its own SSE field (reasoning_content on DeepSeek, Qwen and GLM; reasoning on OpenRouter). Send it to the other entry point:

for await event in client.streamEvents(prompt) {
    switch event {
    case .reasoning(let text): chat.appendReasoningChunk(text, to: "reply")
    case .content(let text):   chat.appendChunk(text, to: "reply")
    }
}

When you hand over whole messages rather than chunks, the same text goes in MarkdownChatMessage.reasoning:

func receive(reasoning text: String, for replyID: String, in messages: inout [MarkdownChatMessage]) {
    guard let index = messages.firstIndex(where: { $0.id == replyID }) else { return }
    messages[index].reasoning += text
}

Either way it lands in a fold above the message: open and streaming while the model thinks, collapsed to a single tappable line the moment the answer starts. A reader who opens or closes it themselves is never overruled afterwards.

Copying a message copies the answer, not the reasoning.

var options = MarkdownChatRenderOptions()
options.reasoning.thinkingLabel = "Thinking…"
options.reasoning.completedLabel = "Thought process"
options.reasoning.durationLabelFormat = "Thought for {seconds}s"
options.reasoning.isExpandedByDefault = false
SettingDefaultEffect
isEnabledtrueTurn off to drop reasoning and leave inline markers as literal text.
parsesInlineTagstrueTurn off when reasoning only ever arrives through appendReasoningChunk.
tagsthink, thinking, thought, reasoning, reasonMarker names recognised in the content stream.
collapsesWhenAnswerBeginstrueFold away as soon as the answer starts.
isExpandedByDefaultfalseWhether a finished message shows its reasoning already open.

finishReasoning(for:) closes the fold early; setReasoningExpanded(_:for:) drives it from your own UI. Both are also on MarkdownChatRenderProxy.

How text appears

options.streamingPresentation = .smooth      // default
options.streamingPresentation = .immediate

.smooth paces the reveal per animation frame, so an answer arriving in uneven network bursts still types out evenly. .immediate puts every chunk on screen the instant it arrives.

Long blocks without a blank line — a wide table, a hundred-line code block — would otherwise be re-parsed on every frame. Two strategies bound that:

options.renderOptions.longBlockStreaming = .throttled   // default: keep moving, repaint less often
options.renderOptions.longBlockStreaming = .deferred    // hold the block, paint it once when complete
options.renderOptions.longBlockThreshold = 512          // characters before either kicks in

Unfinished markdown is handled rather than exposed: a formula whose closing brace has not arrived yet stays in the body colour instead of flashing KaTeX red, and a half-typed code fence is never opened for one frame and closed on the next.

Scrolling

options.autoScrollBehavior = .nearBottom   // default
options.autoScrollBehavior = .always
options.autoScrollBehavior = .never
options.preservesUserScroll = true
options.bottomThreshold = 96

.nearBottom follows new tokens only while the reader is already near the bottom, so scrolling up to re-read is never interrupted.

To drive scrolling yourself — a floating "jump to latest" button, for example — hold a proxy and watch the viewport:

struct ConversationScreen: View {
    @State private var proxy = MarkdownChatRenderProxy()
    @State private var isNearBottom = true
    let messages: [MarkdownChatMessage]

    var body: some View {
        MarkdownChatRenderView(messages: messages, proxy: proxy) { event in
            if case .viewportChanged(let viewport) = event {
                isNearBottom = viewport.isNearBottom
            }
        }
        .overlay(alignment: .bottom) {
            if !isNearBottom {
                Button("Jump to latest") { proxy.scrollToBottom() }
            }
        }
    }
}

Message actions

Copy, retry and edit render as an accessible toolbar under each message.

options.messageActions.assistantActions = [.copy, .retry]
options.messageActions.userActions = [.copy, .edit]

Copy is on by default and the SDK performs it end to end. Retry and edit are not, and the reason is worth stating plainly: the SDK cannot re-run your request or write into your composer. Listing an action is your promise to handle it — a button that does nothing looks finished and is not.

MarkdownChatRenderView(messages: messages, options: options) { event in
    guard case .messageAction(let action) = event else { return }
    switch action.action {
    case .retry: regenerate(messageID: action.messageID)
    case .edit:  moveToComposer(messageID: action.messageID)
    case .copy:  break   // already done; this is just the receipt
    @unknown default: break
    }
}

In a debug build, emitting an action nobody handles prints a line to the console explaining what was ignored.