protobuf

Commit Graph

Author	SHA1	Message	Date
Eric Salo	3475ebec94	upb: move the split64 accessors out of upbc/ PiperOrigin-RevId: 527396510	2 years ago
Joshua Haberman	55cfaf3e0c	Moved promotion-related accessors to a separate file. Since promotion is a more complicated operation than the simple accessors, and since promotion logic will likely be changing before long, it helps to put promotion-related logic in a separate place and rule. PiperOrigin-RevId: 525519707	2 years ago
Joshua Haberman	b81ae408b0	Fixed the upb.lexan build: - Fixed a couple of broken tests that were probably invoking UB. - Excluded python/... and js/..., as these do not work with Windows. PiperOrigin-RevId: 525228589	2 years ago
Protobuf Team Bot	433e737c0e	[C++ Upb] Implement copy constructor and assignment operator. PiperOrigin-RevId: 524941810	2 years ago
Joshua Haberman	a49ff5513e	Split out `message/accessors_internal.h` and added `upb_Arena` assertions. These changes will help clarify and harden the `accessors.h` interface. PiperOrigin-RevId: 524856618	2 years ago
Joshua Haberman	df93cf65a2	Hide upb_MiniTableField.submsg_index with new `UPB_PRIVATE()` macro The fields of upb_MiniTableField are intended to be internal-only, accessed only through public functions like `upb_MiniTable_GetSubMessageTable()`. But over time, clients have started accessing many of these fields directly. This is an easy mistake to make, as there is no clear signal that the fields should not be used in applications. This makes the implementation difficult to change without breaking users. The new `UPB_PRIVATE()` macro appends an unpredictable string to each private symbol. This makes it very difficult to accidentally use a private symbol, since users would need to write something like `field->submsg_index_dont_copy_me__upb_internal_use_only`. This is still possible to do, but it leaves a clear wart in the code showing that an an encapsulation break has occurred. The `UPB_PRIVATE()` macro itself is defined in `port/def.inc`, which users cannot include directly. Once we land this, more such CLs will follow for the other fields of `upb_MiniTable*`. We will add inline functions as needed to provide the semantic functionality needed by users. PiperOrigin-RevId: 523166901	2 years ago
Protobuf Team Bot	9d2b5d1716	Add upb_MiniTable_FindUnknown depth_limit parameter. Fix UpbMessageIsEqual IsMap fallthrough. PiperOrigin-RevId: 522961305	2 years ago
Mike Kruskal	d260ab343e	Add windows CI PiperOrigin-RevId: 520478558	2 years ago
Joshua Haberman	6d6ab90ece	Enable arena test (it was previously #ifdef'd away accidentally) PiperOrigin-RevId: 519782179	2 years ago
Joshua Haberman	d450990631	Allow for fuse/free races in `upb_Arena`. Implementation is by kfm@, I only added the portability code around it. `upb_Arena` was designed to be only thread-compatible. However, fusing of arenas muddies the waters somewhat, because two distinct `upb_Arena` objects will end up sharing state when fused. This causes a `upb_Arena_Free(a)` to interfere with `upb_Arena_Fuse(b, c)` if `a` and `b` were previously fused. It turns out that we can use atomics to fix this with about a 35% regression in fuse performance (see below). Arena create+free does not regress, thanks to special-case logic in Free(). `upb_Arena` is still a thread-compatible type, and it is still never safe to call `upb_Arena_xxx(a)` and `upb_Arena_yyy(a)` in parallel. However you can at least now call `upb_Arena_Free(a)` and `upb_Arena_Fuse(b, c)` in parallel, even if `a` and `b` were previously fused. Note that `upb_Arena_Fuse(a, b)` and `upb_Arena_Fuse(c, d)` is still not allowed if `b` and `c` were previously fused. In practice this means that fuses must still be single-threaded within a single fused group. Performance results: ``` name old cpu/op new cpu/op delta BM_ArenaOneAlloc 18.6ns ± 1% 18.6ns ± 1% ~ (p=0.726 n=18+17) BM_ArenaInitialBlockOneAlloc 6.28ns ± 1% 5.73ns ± 1% -8.68% (p=0.000 n=17+20) BM_ArenaFuseUnbalanced/2 44.1ns ± 2% 60.4ns ± 1% +37.05% (p=0.000 n=18+19) BM_ArenaFuseUnbalanced/8 370ns ± 2% 500ns ± 1% +35.12% (p=0.000 n=19+20) BM_ArenaFuseUnbalanced/64 3.52µs ± 1% 4.71µs ± 1% +33.80% (p=0.000 n=18+19) BM_ArenaFuseUnbalanced/128 7.20µs ± 1% 9.72µs ± 2% +34.93% (p=0.000 n=16+19) BM_ArenaFuseBalanced/2 44.4ns ± 2% 61.4ns ± 1% +38.23% (p=0.000 n=20+17) BM_ArenaFuseBalanced/8 373ns ± 2% 509ns ± 1% +36.57% (p=0.000 n=19+17) BM_ArenaFuseBalanced/64 3.55µs ± 2% 4.79µs ± 1% +34.80% (p=0.000 n=19+19) BM_ArenaFuseBalanced/128 7.26µs ± 1% 9.76µs ± 1% +34.45% (p=0.000 n=17+19) BM_LoadAdsDescriptor_Upb<NoLayout> 5.66ms ± 1% 5.69ms ± 1% +0.57% (p=0.013 n=18+20) BM_LoadAdsDescriptor_Upb<WithLayout> 6.30ms ± 1% 6.36ms ± 1% +0.90% (p=0.000 n=19+18) BM_LoadAdsDescriptor_Proto2<NoLayout> 12.1ms ± 1% 12.1ms ± 1% ~ (p=0.118 n=18+18) BM_LoadAdsDescriptor_Proto2<WithLayout> 12.2ms ± 1% 12.3ms ± 1% +0.50% (p=0.006 n=18+18) BM_Parse_Upb_FileDesc<UseArena, Copy> 12.7µs ± 1% 12.7µs ± 1% ~ (p=0.194 n=20+19) BM_Parse_Upb_FileDesc<UseArena, Alias> 11.6µs ± 1% 11.6µs ± 1% ~ (p=0.192 n=20+20) BM_Parse_Upb_FileDesc<InitBlock, Copy> 12.5µs ± 1% 12.5µs ± 0% ~ (p=0.750 n=18+14) BM_Parse_Upb_FileDesc<InitBlock, Alias> 11.4µs ± 1% 11.3µs ± 1% -0.34% (p=0.046 n=19+19) BM_Parse_Proto2<FileDesc, NoArena, Copy> 25.4µs ± 1% 25.7µs ± 2% +1.37% (p=0.000 n=18+18) BM_Parse_Proto2<FileDesc, UseArena, Copy> 12.1µs ± 2% 12.1µs ± 1% ~ (p=0.143 n=18+18) BM_Parse_Proto2<FileDesc, InitBlock, Copy> 11.9µs ± 3% 11.9µs ± 1% ~ (p=0.076 n=17+19) BM_Parse_Proto2<FileDescSV, InitBlock, Alias> 13.2µs ± 1% 13.2µs ± 1% ~ (p=0.053 n=19+19) BM_SerializeDescriptor_Proto2 5.97µs ± 4% 5.90µs ± 4% ~ (p=0.093 n=17+19) BM_SerializeDescriptor_Upb 10.4µs ± 1% 10.4µs ± 1% ~ (p=0.909 n=17+18) name old time/op new time/op delta BM_ArenaOneAlloc 18.7ns ± 2% 18.6ns ± 0% ~ (p=0.607 n=18+17) BM_ArenaInitialBlockOneAlloc 6.29ns ± 1% 5.74ns ± 1% -8.71% (p=0.000 n=17+19) BM_ArenaFuseUnbalanced/2 44.1ns ± 1% 60.6ns ± 1% +37.21% (p=0.000 n=17+19) BM_ArenaFuseUnbalanced/8 371ns ± 2% 500ns ± 1% +35.02% (p=0.000 n=19+16) BM_ArenaFuseUnbalanced/64 3.53µs ± 1% 4.72µs ± 1% +33.85% (p=0.000 n=18+19) BM_ArenaFuseUnbalanced/128 7.22µs ± 1% 9.73µs ± 2% +34.87% (p=0.000 n=16+19) BM_ArenaFuseBalanced/2 44.5ns ± 2% 61.5ns ± 1% +38.22% (p=0.000 n=20+17) BM_ArenaFuseBalanced/8 373ns ± 2% 510ns ± 1% +36.58% (p=0.000 n=19+16) BM_ArenaFuseBalanced/64 3.56µs ± 2% 4.80µs ± 1% +34.87% (p=0.000 n=19+19) BM_ArenaFuseBalanced/128 7.27µs ± 1% 9.77µs ± 1% +34.40% (p=0.000 n=17+19) BM_LoadAdsDescriptor_Upb<NoLayout> 5.67ms ± 1% 5.71ms ± 1% +0.60% (p=0.011 n=18+20) BM_LoadAdsDescriptor_Upb<WithLayout> 6.32ms ± 1% 6.37ms ± 1% +0.87% (p=0.000 n=19+18) BM_LoadAdsDescriptor_Proto2<NoLayout> 12.1ms ± 1% 12.2ms ± 1% ~ (p=0.126 n=18+19) BM_LoadAdsDescriptor_Proto2<WithLayout> 12.2ms ± 1% 12.3ms ± 1% +0.51% (p=0.002 n=18+18) BM_Parse_Upb_FileDesc<UseArena, Copy> 12.7µs ± 1% 12.7µs ± 1% ~ (p=0.149 n=20+19) BM_Parse_Upb_FileDesc<UseArena, Alias> 11.6µs ± 1% 11.6µs ± 1% ~ (p=0.211 n=20+20) BM_Parse_Upb_FileDesc<InitBlock, Copy> 12.5µs ± 1% 12.5µs ± 1% ~ (p=0.986 n=18+15) BM_Parse_Upb_FileDesc<InitBlock, Alias> 11.4µs ± 1% 11.3µs ± 1% ~ (p=0.081 n=19+18) BM_Parse_Proto2<FileDesc, NoArena, Copy> 25.4µs ± 1% 25.8µs ± 2% +1.41% (p=0.000 n=18+18) BM_Parse_Proto2<FileDesc, UseArena, Copy> 12.1µs ± 2% 12.1µs ± 1% ~ (p=0.558 n=19+18) BM_Parse_Proto2<FileDesc, InitBlock, Copy> 12.0µs ± 3% 11.9µs ± 1% ~ (p=0.165 n=17+19) BM_Parse_Proto2<FileDescSV, InitBlock, Alias> 13.2µs ± 1% 13.2µs ± 1% ~ (p=0.070 n=19+19) BM_SerializeDescriptor_Proto2 5.98µs ± 4% 5.92µs ± 3% ~ (p=0.138 n=17+19) BM_SerializeDescriptor_Upb 10.4µs ± 1% 10.4µs ± 1% ~ (p=0.858 n=17+18) ``` PiperOrigin-RevId: 518573683	2 years ago
Matt Kulukundis	5bc3cae2d7	Add threading tests for arenas PiperOrigin-RevId: 518009578	2 years ago
Protobuf Team Bot	3286f948f8	Implements upb_Message_DeepClone. PiperOrigin-RevId: 514723111	2 years ago
Joshua Haberman	57a79de7cc	Ensure that extensions respect deterministic serialization. Previously we were not sorting extensions at encode time, even in deterministic mode. PiperOrigin-RevId: 508217926	2 years ago
Eric Salo	4843fd0d75	move conformance tests into a separate subdir PiperOrigin-RevId: 502744594	2 years ago
Joshua Haberman	e41a2d7ba0	upb is self-hosting! This CL changes the upb compiler to no longer depend on C++ protobuf libraries. upb now uses its own reflection libraries to implement its code generator. # Key Benefits 1. upb can now use its own reflection libraries throughout the compiler. This makes upb more consistent and principled, and gives us more chances to dogfood our own C++ reflection API. This highlighted several parts of the C++ reflection API that were incomplete. 2. This CL removes code duplication that previously existed in the compiler. The upb reflection library has code to build MiniDescriptors and MiniTables out of descriptors, but prior to this CL the upb compiler could not use it. The upb compiler had a separate copy of this logic, and the compiler's copy of this logic was especially tricky and hard to maintain. This CL removes the separate copy of that logic. 3. This CL (mostly) removes upb's dependency on the C++ protobuf library. We still depend on `protoc` (the binary), but the runtime and compiler no longer link against C++'s libraries. This opens up the possibility of speeding up some builds significantly if we can use a prebuilt `protoc` binary. # Bootstrap Stages To bootstrap, we check in a copy of our generated code for `descriptor.proto` and `plugin.proto`. This allows the compiler to depend on the generated code for these two protos without creating a circular dependency. This code is checked in to the `stage0` directory. The bootstrapping process is divided into a few stages. All `cc_library()`, `upb_proto_library()`, and `cc_binary()` targets that would otherwise be circular participate in this staging process. That currently includes: * `//third_party/upb:descriptor_upb_proto` * `//third_party/upb:plugin_upb_proto` * `//third_party/upb:reflection` * `//third_party/upb:reflection_internal` * `//third_party/upbc:common` * `//third_party/upbc:file_layout` * `//third_party/upbc:plugin` * `//third_party/upbc:protoc-gen-upb` For each of these targets, we produce a rule for each stage (the logic for this is nicely encapsulated in Blaze/Bazel macros like `bootstrap_cc_library()` and `bootstrap_upb_proto_library()`, so the `BUILD` file remains readable). For example: * `//third_party/upb:descriptor_upb_proto_stage0` * `//third_party/upb:descriptor_upb_proto_stage1` * `//third_party/upb:descriptor_upb_proto` The stages are: 1. `stage0`: This uses the checked-in version of the generated code. The stage0 compiler is correct and outputs the same code as all other compilers, but it is unnecessarily slow because its protos were compiled in bootstrap mode. The stage0 compiler is used to generate protos for stage1. 2. `stage1`: The stage1 compiler is correct and fast, and therefore we use it in almost all cases (eg. `upb_proto_library()`). However its own protos were not generated using `upb_proto_library()`, so its `cc_library()` targets cannot be safely mixed with `upb_proto_library()`, as this would lead to duplicate symbols. 3. final (no stage): The final compiler is identical to the `stage1` compiler. The only difference is that its protos were built with `upb_proto_library()`. This doesn't matter very much for the compiler binary, but for the `cc_library()` targets like `//third_party/upb:reflection`, only the final targets can be safely linked in by other applications. # "Bootstrap Mode" Protos The checked-in generated code is generated in a special "bootstrap" mode that is a bit different than normal generated code. Bootstrap mode avoids depending on the internal representation of MiniTables or the messages, at the cost of slower runtime performance. Bootstrap mode only interacts with MiniTables and messages using public APIs such as `upb_MiniTable_Build()`, `upb_Message_GetInt32()`, etc. This is very important as it allows us to change the internal representation without needing to regenerate our bootstrap protos. This will make it far easier to write CLs that change the internal representation, because it avoids the awkward dance of trying to regenerate the bootstrap protos when the compiler itself is broken due to bootstrap protos being out of date. The bootstrap generated code does have two downsides: 1. The accessors are less efficient, because they look up MiniTable fields by number instead of hard-coding the MiniTableField into the generated code. 2. It requires runtime initialization of the MiniTables, which costs CPU cycles at startup, and also allocates memory which is never freed. Per google3 rules this is not really a leak, since this memory is still reachable via static variables, but it is undesirable in many contexts. We could fix this part by introducing the equivalent of `google::protobuf::ShutdownProtobufLibrary()`). These downsides are fine for the bootstrapping process, but they are reason enough not to enable bootstrap mode in general for all protos. # Bootstrapping Always Uses OSS Protos To enable smooth syncing between Google3 and OSS, we always use an OSS version of the checked in generated code for `stage0`, even in google3. This requires that the google3 code can be switched to reference the OSS proto names using a preprocessor define. We introduce the `UPB_DESC(xyz)` macro for this, which will expand into either `proto2_xyz` or `google_protobuf_xyz`. Any libraries used in `stage0` must use `UPB_DESC(xyz)` rather than refer to the symbol names directly. PiperOrigin-RevId: 501458451	2 years ago
Joshua Haberman	143132fa27	Make upb's generated code agnostic to fasttable. This simplifies the code generation by making output agnostic to whether fasttables will be used or not. This grows the generated code in the common case, but when fasttables are not being used the preprocessor will strip away the unused tables. PiperOrigin-RevId: 499340805	2 years ago
Deanna Garcia	f887fe30aa	Add lots more source files	2 years ago
Joshua Haberman	7cd8f6c940	Ported more cases of wire format parsing to upb_WireReader. PiperOrigin-RevId: 498502557	2 years ago
Joshua Haberman	112f037da6	Migrated the text format encoder to upb_WireReader. Also fixed a few bugs and added new functionality to upb_WireReader. PiperOrigin-RevId: 498223970	2 years ago
Eric Salo	26e9a75294	remove wire/types.h from the :wire build target PiperOrigin-RevId: 498167993	2 years ago
Joshua Haberman	75488a0742	Created a new upb_WireReader interface for parsing wire data directly. The overall motivation for this interface is to consolidate many places in upb that are parsing wire format data directly. This interface is not yet complete, but this is a good start. We have enough to port the wire format parsing in accessors.c to this interface. We can follow up by porting more places that do wire format parsing. PiperOrigin-RevId: 498109788	2 years ago
Joshua Haberman	a48af3f824	Moved aliasing logic for string field parsing into EpsCopyInputStream. Moving the logic down to EpsCopyInputStream makes it easier to test and reuse this functionality. We also implement aliasing for the final bytes of the patch buffer, which has never been supported before. We used to always force a copy for any data parsed out of the patch buffer at the end of the stream. Much of this logic is ported directly from the C++ EpsCopyInputStream class. PiperOrigin-RevId: 498091644	2 years ago
Eric Salo	02cf7aaa1d	upb_Array_Resize() now correctly clears new values PiperOrigin-RevId: 496255949	2 years ago
Joshua Haberman	68d1d91475	Separated out buffering code into upb_EpsCopyInputStream. This mirrors the structure of C++ protobuf, which has an EpsCopyInputStream class. This will lay the foundation for making EpsCopyInputStream capable of true streaming, by reading its input from a ZeroCopyInputStream. It also lets us test EpsCopyInputStream separately from the decoder: see the new unit test that fuzzes upb_EpsCopyInputStream. After this CL is submitted, the two decoders (the normal decoder and the fast decoder) should no longer be accessing the members of upb_EpsCopyInputStream. PiperOrigin-RevId: 494400285	2 years ago
Mike Kruskal	1fee6d8327	Expand visibility of amalgamation targets. This will give us more freedom to refactor PHP and Ruby runtimes as part of go/protobuf-bazelification. PiperOrigin-RevId: 493074147	2 years ago
Eric Salo	b98c7c7b50	move non-conformance tests and protos into upb/test/ create upb/test/BUILD PiperOrigin-RevId: 492534429	2 years ago
Mike Kruskal	3e078f5fe4	Add CMake+Bazel dependencies on utf8_range repo	2 years ago
Eric Salo	a4a7c30b48	remove an obsolete dependency on :message from :collections PiperOrigin-RevId: 491197910	2 years ago
Eric Salo	4f098fd54c	move the extension registry down into mini_table/ Now that :minitable depends on :hash, we may as well put the extreg there also. PiperOrigin-RevId: 490620348	2 years ago
Joshua Haberman	9c6223c058	Unified all hazzers to use MiniTable accessors This required some work to unify map entry messages with regular messages, with respect to presence. Before map entry fields could never have presence. Now they can have presence according to normal rules. Note that this only applies to times that the user constructs a map entry directly. PiperOrigin-RevId: 490611656	2 years ago
Eric Salo	ff6439fba0	move the wire type definitions into upb/wire/ where they belong PiperOrigin-RevId: 489500430	2 years ago
Eric Salo	384ffc0af8	implement reserved names and ranges for messages and enums https://github.com/protocolbuffers/protobuf/issues/10158 PiperOrigin-RevId: 489285657	2 years ago
Eric Salo	fb7a67458f	add build targets for :wire and :message and :message_internal PiperOrigin-RevId: 489267642	2 years ago
Eric Salo	27d70edfe2	clean up the :mini_table build target Remove circular dependencies that were bouncing back and forth between msg_internal.h and mini_table/, including: - splitting out each mini table subtype into its own header - moving the non-reflection message code into message/ - moving the accessors from mini_table/ to message/ PiperOrigin-RevId: 489121042	2 years ago
Laramie Leavitt	7d76251d9f	Split upb_amalgamation into a separate .bzl file PiperOrigin-RevId: 489034477	2 years ago
Eric Salo	4d3998b54b	consolidate some general parsing functions into upb/lex/ There are several other functions which might eventually end up here and ideally become unified across json/ and text/ and io/ so this is just a first step to create the new subdir and get rid of upb/internal/ PiperOrigin-RevId: 488954926	2 years ago
Joshua Haberman	023c4da591	Enabled TAP testing for upb on Windows via Lexan. We disable targets that are not currently working on Windows. PiperOrigin-RevId: 488560033	2 years ago
Eric Salo	b3cb3fbea8	create upb/hash/ The next lowest build target to scrub is the hash table. We already have a few other things called 'table' (mini table, fast table) so let's just go with 'hash' here. Split apart the headers into int and str branches sharing common definitions. Leave the core functions in a single .c for inlining. PiperOrigin-RevId: 488388767	2 years ago
Eric Salo	ff8e1b40ba	create base/ subdir and expand :status build target to :base upb.h is now just a temporary stub PiperOrigin-RevId: 488255988	2 years ago
Eric Salo	632471333c	clean up the build targets for collections, mem, reflection PiperOrigin-RevId: 488226166	2 years ago
Eric Salo	33633fd604	generated code now uses the scalar get accessors Also tweaked the generator to only emit a call to UPB_SIZE() when the two input values are actually different. PiperOrigin-RevId: 487606926	2 years ago
Eric Salo	aec12a466f	upb: split out :status as a separate build target This should allow other upb components to depend upon the zcis without causing a cycle PiperOrigin-RevId: 486987806	2 years ago
Eric Salo	a77b9665e1	move lua/ up to the top level directory where python/ lives PiperOrigin-RevId: 486786325	2 years ago
Eric Salo	f6307877d3	move portability stuff into upb/port/ Also delete redundant system #includes that are already pulled in by port/def.inc PiperOrigin-RevId: 486398989	2 years ago
Eric Salo	46699b72ad	move message set enums into upb/wire/ (and use them) PiperOrigin-RevId: 486366363	2 years ago
Eric Salo	fd040a8bff	create collections/map_internal.h and collections/map_gencode_util.h Move the map-related functions from msg_internal.h that are only used in generated code into map_gencode_util.h. Then move the rest of the map-related functions from msg_internal.h into map_internal.h. PiperOrigin-RevId: 486299140	2 years ago
Protobuf Team Bot	a4779ef5f8	internal change PiperOrigin-RevId: 486275169	2 years ago
Protobuf Team Bot	0f4fffef16	Update config_setting visibility in support of --incompatible_config_setting_private_default_visibility. For https://github.com/bazelbuild/bazel/issues/12933. PiperOrigin-RevId: 486230747	2 years ago
Eric Salo	fd14316f38	create collections/ subdir for all array and map code PiperOrigin-RevId: 486013554	2 years ago
Eric Salo	d9b6f13cde	remove upb_MtDataEncoder from the public surface PiperOrigin-RevId: 485928803	2 years ago

1 2 3 4 5 ...

309 Commits (3475ebec9484eb61e35c43c1aace669c4d16a25f)