Skip to content

TCP N9: POCO row insert — InsertAsync<T> over a compiled per-column gather - #559

Draft
alex-clickhouse wants to merge 3 commits into
tcp/epic-n5-poco-readfrom
tcp/epic-n9-poco-write
Draft

TCP N9: POCO row insert — InsertAsync<T> over a compiled per-column gather#559
alex-clickhouse wants to merge 3 commits into
tcp/epic-n5-poco-readfrom
tcp/epic-n9-poco-write

Conversation

@alex-clickhouse

Copy link
Copy Markdown
Collaborator

Stacked on #557 (tcp/epic-n5-poco-read) — review that one first; this PR's diff is the write half only.

TCP epic N9: row-oriented insert, both shapes. The read half of Branch 2 landed in #557; this is the mirror, so a POCO now round-trips through the native protocol.

await client.InsertAsync("INSERT INTO t (id, user_name) VALUES", accounts);      // POCO rows
await client.InsertAsync("INSERT INTO t (id, user_name) VALUES", objectRows);    // untyped, positional

What it does

InsertAsync<T>(IEnumerable<T>) — one compiled gather per target column, each pulling one property out of every row into the buffer that column is written from. Box-free (except a Variant/Dynamic target, written from object), no per-row delegate hop, and for a fixed-width column the gathered buffer reaches the wire as a single blit. Names match as the read side's do: case-, then underscore-insensitive, with [ClickHouseTcpColumn] / [ClickHouseTcpNotMapped].

InsertAsync(IEnumerable<object[]>) — the dynamic tier, positional by the sample block's order. Each column's CLR type comes from the first row that has a value there, so hand-written DateTime values and the raw epoch seconds the untyped read produces are both accepted, and a read-then-reinsert needs no conversion by the caller.

No conversion layer of its own. The read side must ask the codec for its conversions, because a column decodes to the raw wire value. Write does not: a codec already accepts DateTime/DateTimeOffset/TimeSpan directly, so the only conversions left are the CLR-level ones a cast would do — nullable lift, enum ordinal, reference upcast. Numeric widening is declined in both directions, which is what keeps "inserts" and "reads back" the same set of shapes.

Decisions worth a look

  • Every target column must map to a property, unlike a read, where an unmapped column is skipped: the server expects a value for each column of the statement's list. The remedy the message names is the statement's own column list, which is also how one POCO fills part of a table. A property with no column stays silent.
  • A mapping error lands mid-INSERT and must not cost the connection. The target types only arrive in the sample block, so the columns are built there, behind a new internal InsertColumnFactory seam. A factory that throws parks its exception, closes the row stream with no rows, drains, and rethrows once the connection is back to Ready — the course the schema-mismatch path already took. The columns a factory returns are the insert's to dispose; a caller's own columns are untouched, as before.
  • "Can this column hold a null" is the codec's question. NullPlaceholder is null is true exactly for Nullable, a nullable LowCardinality, Variant and Dynamic. Asking the CLR instead (default(TWrite) is null) let a null string property into a plain String column through the gather, to fault inside WriteColumn part-way through a block — which terminates the connection. Caught in review; the gather now compiles a null test for reference-typed properties too, so it fails before anything is sent, naming the row.
  • Nullable's WritableElementTypes was under-reporting. CanWrite has always accepted DateTime? for a Nullable(DateTime) column, but the list reported only the canonical uint? — invisible to a caller who probes with a column, fatal to one that picks a type from the list. Now lifted, along with NullPlaceholderAs.
  • The write plan's cache key omits the session timezone the read key carries: the write path resolves in ResolveContext.ForWrite, so one target shape is one plan. Both plans now share PocoBlockSignature for the injective length-prefixed key.

Two things to weigh

  1. Overload resolution. Adding InsertAsync<T> makes InsertAsync(sql, null) and a generically typed column sequence (IColumn<ulong>[]) ambiguous — IReadOnlyList<IColumn> and IEnumerable<IColumn<ulong>> are unrelated, so the "non-generic beats generic" tie-break never runs. A plain IColumn[]/List<IColumn> still binds to the columnar overload, which is every call site here and the only shape an external caller can build, the concrete column classes being internal. Kept because it mirrors the shipped QueryAsync/QueryAsync<T> pairing and the assembly is unreleased and [Experimental]; InsertRowsAsync is the escape hatch if that changes. An IEnumerable<IColumn> (a LINQ operator over a column list) now binds to the POCO overload and gets a guard naming the columnar one.
  2. LowCardinality(DateTime) reads as DateTime but is written only from uint — it lifts its inner's readable types and not its writable ones, so it is the last codec whose two surfaces disagree. Pre-existing, and the columnar path has always had it, but the POCO layer is what makes it visible. Documented in PocoWriteConversion and left as a follow-up, since closing it means giving the LowCardinality write path a shape per write type as Nullable has.

Testing

2011 tests pass on net9 against a real server; every new Poco/ file is at 100% line coverage.

  • Per-type coverage rides the corpus: all 203 InsertRoundTripCase cases are inserted as Row<T> and read back. Nested is the one shape rows cannot fill — its writer needs its own column type — so the corpus test asserts the refusal for it rather than skipping.
  • POCO→POCO round trip, calendar and enum properties (including the nullable spellings), an insert-only immutable POCO, a narrower column list, a lazy 5 000-row source across several blocks, and the untyped path's positional round trip and read→re-insert.
  • The mid-INSERT refusals all assert the client is still usable afterwards, and the connection-level tests pin that the connection itself stays Ready rather than being redialled, plus the factory's ownership (dispose count) and that the factory's own exception comes back with its identity intact.
  • Unit tests cover only what a round trip cannot see: which CLR type the plan chose, the plan-build refusals, and the cache key (separator-spelling collision, same-name-different-type, timezone independence).

No CHANGELOG/RELEASENOTES entry: no TCP epic PR has one, the assembly being unreleased and [Experimental]. The epic gets a single entry when it ships.

🤖 Generated with Claude Code

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the TCP client’s row-oriented insert half, complementing #557’s POCO query support.

Changes:

  • Adds POCO and positional object[] insert overloads.
  • Compiles cached per-column gather plans with codec-aware null/type handling.
  • Preserves connections across mapping failures and adds broad unit/integration coverage.

Reviewed changes

Copilot reviewed 21 out of 21 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
ClickHouse.Driver.Tcp/Types/Codecs/NullableColumnCodec.cs Exposes lifted writable types.
ClickHouse.Driver.Tcp/Types/ArrayColumn.cs Adds pooled-buffer ownership.
ClickHouse.Driver.Tcp/Protocol/ClickHouseTcpConnection.cs Adds schema-driven column factories.
ClickHouse.Driver.Tcp/Poco/PocoWritePlan.cs Builds compiled POCO write plans.
ClickHouse.Driver.Tcp/Poco/PocoWriteConversion.cs Resolves property-to-codec conversions.
ClickHouse.Driver.Tcp/Poco/PocoUntypedColumns.cs Transposes positional rows.
ClickHouse.Driver.Tcp/Poco/PocoTypeRegistry.cs Caches write plans.
ClickHouse.Driver.Tcp/Poco/PocoTypeDescriptor.cs Shares mapped-column descriptions.
ClickHouse.Driver.Tcp/Poco/PocoRowBuffer.cs Materializes row sources.
ClickHouse.Driver.Tcp/Poco/PocoReadPlan.cs Uses shared block signatures.
ClickHouse.Driver.Tcp/Poco/PocoColumnBuilder.cs Compiles per-column gathers.
ClickHouse.Driver.Tcp/Poco/PocoBlockSignature.cs Centralizes plan cache keys.
ClickHouse.Driver.Tcp/Client/IClickHouseTcpClient.cs Adds row-insert contracts.
ClickHouse.Driver.Tcp/Client/ClickHouseTcpClient.cs Implements row inserts.
ClickHouse.Driver.Tcp.Tests/Types/NullableColumnCodecTests.cs Tests lifted nullable writes.
ClickHouse.Driver.Tcp.Tests/Types/ArrayColumnTests.cs Tests buffer ownership.
ClickHouse.Driver.Tcp.Tests/Protocol/ClickHouseTcpConnectionInsertTests.cs Tests factory lifecycle.
ClickHouse.Driver.Tcp.Tests/Poco/PocoWritePlanTests.cs Tests write-plan behavior.
ClickHouse.Driver.Tcp.Tests/Poco/PocoUntypedColumnsTests.cs Tests positional transposition.
ClickHouse.Driver.Tcp.Tests/Poco/PocoRowBufferTests.cs Tests row materialization.
ClickHouse.Driver.Tcp.Tests/Integration/PocoWriteIntegrationTests.cs Exercises real-server round trips.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +64 to +65
foreach (T row in source)
{

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct, and fixed in aa3978d — the check is now per row rather than at the growth points, so a counted source (which rents once to fit and never grows) is no longer drained in full. The read is a field test against a token that is usually None, so it costs nothing measurable beside the source’s own MoveNext.

Added Materialize_TokenCancelledPartWayThroughACountedSource_StopsThere, which needs a source that reports its Count (so the counted path is taken) and cancels the token part-way through yielding — a plain List<T> cannot express that.

@codecov

codecov Bot commented Aug 18, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

alex-clickhouse and others added 3 commits August 18, 2026 10:01
The write half of Branch 2 (N9), stacked on the read half. A row-oriented
insert transposes the caller's rows into the columnar insert the connection
already sends, in both shapes: `InsertAsync<T>(IEnumerable<T>)` gathering each
property into the buffer its target column is written from, and
`InsertAsync(IEnumerable<object[]>)` matching values to targets by position.

The gather is the mirror of the read scatter: one compiled loop per target
column, no boxing and no per-row delegate hop. It needs no conversion layer of
its own — a codec already accepts the calendar types on write, so the only
conversions left are the CLR-level ones a cast would do (nullable lift, enum
ordinal, reference upcast). Numeric widening is declined in both directions, so
every shape that inserts also reads back.

Whether a row may have no value for a column is the codec's question, not the
CLR's: `NullPlaceholder is null` is true exactly for the types with a NULL of
their own, so a null `string` property into a plain `String` column is reported
before anything is sent, naming the row, rather than faulting inside the codec
part-way through a block and taking the connection with it.

The target types arrive in the server's sample block, after the statement has
gone out, so the connection gains an InsertColumnFactory seam that builds the
columns there and owns them afterwards. A factory that throws — a mapping
error, which is the caller's shape rather than the connection's — closes the
row stream with no rows and reports once the connection is back to Ready,
exactly as a schema mismatch does.

Also lifts `Nullable`'s WritableElementTypes to its nullable surface. CanWrite
has always accepted `DateTime?` for a `Nullable(DateTime)` column, but the list
reported only the canonical `uint?`, which a plan choosing a write type from
the list rather than probing with a column cannot see.

Co-Authored-By: Claude <noreply@anthropic.com>
The check sat at the buffer's growth points, which a counted source never
reaches — it rents once to fit — so a long `List<T>` was drained in full
whatever the token said. Tested per row instead, next to the null-row check:
the read is a field test against a token that is usually None, so it costs
nothing measurable beside the source's own MoveNext.

Co-Authored-By: Claude <noreply@anthropic.com>
The corpus insert test knows a Nested target cannot be gathered from rows and
asserts the refusal instead of a successful insert. It recognised the shape
with StartsWith, so it only caught a top-level Nested -- and the corpus also
has Array(Nested(a UInt8)), Tuple(Nested(a UInt8), String) and Nested in both
Map positions. Those four expected an insert that cannot work, and failed on
every framework and every server version.

Contains, not StartsWith: a composite can only hand its child the column shape
a row yields, so a Nested inside one is exactly as ungatherable, and refuses
for the same reason with the same message.

This mirrors b75f932 on tcp/epic-b9-tls, which made the codec itself refuse a
Nested inside a composite rather than only a top-level one. That commit is on a
different epic line and never reached this branch, so the test kept the narrow
check.

Co-Authored-By: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants