Caulobase

Data

Datasets & resources

Where the primary data lives. Every accession below was opened and its title and organism checked against the paper it belongs to (last checked 2026-10-04). A resource that has gone offline is marked as such instead of silently dropped.

On names: NCBI Taxonomy and LPSN now give the correct species name as Caulobacter vibrioides, with Caulobacter crescentus, the name used in nearly all of the research literature, listed as a synonym. The laboratory strain CB15 (ATCC 19089) and its synchronizable derivative NA1000 are not the type strain of either name. source

Genomes & annotation

  1. NA1000 complete genome (RefSeq chromosome) NA1000

    Single circular chromosome, 4,042,929 bp. RefSeq NC_011916.1 = GenBank CP001340.1. Curated RefSeq annotation with CCNA_ locus tags (e.g. ctrA = CCNA_03130).

    NCBI RefSeq NC_011916.1Marks ME et al. 2010 J Bacteriol

  2. NA1000 complete genome (GenBank/INSDC) NA1000

    Original INSDC submission of the NA1000 chromosome from Marks et al. 2010 (University of Chicago annotation, CCNA_ locus tags).

    GenBank CP001340.1Marks ME et al. 2010 J Bacteriol

  3. NA1000 genome assembly ASM2200v1 NA1000

    RefSeq reference genome for the species. GCF_000022005.1 paired with GCA_000022005.1; 1 chromosome, 4.04 Mb; RefSeq annotation 3,886 protein-coding genes (released 2020-12-10).

    NCBI Datasets GCF_000022005.1Marks ME et al. 2010 J Bacteriol

  4. CB15 complete genome (RefSeq chromosome) CB15

    Single circular chromosome, 4,016,947 bp. RefSeq NC_002696.2 = GenBank AE005673.1. RefSeq re-annotated by PGAP (2025-04-26) with CC_RS locus tags; original CC_ tags kept as old locus tags.

    NCBI RefSeq NC_002696.2Nierman WC et al. 2001 PNAS

  5. CB15 complete genome (GenBank/INSDC, TIGR) CB15

    Original TIGR genome submission from Nierman et al. 2001, with the classic CC_#### locus tags used in most pre-2010 literature.

    GenBank AE005673.1Nierman WC et al. 2001 PNAS

  6. CB15 genome assembly ASM690v1 CB15

    GCF_000006905.1 paired with GCA_000006905.1; 1 chromosome, 4.02 Mb. Current RefSeq annotation GCF_000006905.1-RS_2025_04_26 (PGAP 6.10): 3,823 protein-coding genes, 120 pseudogenes.

    NCBI Datasets GCF_000006905.1Nierman WC et al. 2001 PNAS

  7. Type strain DSM 9893 (CB51) draft genome DSM 9893 (CB51)

    Draft (25 contigs, 3.97 Mb) genome of the Caulobacter vibrioides type strain; NCBI flags it 'assembly from type material'. Not the CB15 lineage.

    NCBI Datasets GCF_002858865.1

  8. Strain CB2 complete genome (former C. crescentus type strain) CB2

    Complete 4.12 Mb genome of CB2 (ATCC 15252 / DSM 4727), type strain of the synonym C. crescentus; NCBI flags it 'assembly from heterotypic synonym type material'.

    NCBI Datasets GCF_002310295.2

Databases & portals

  1. NCBI Taxonomy: Caulobacter vibrioides (species) species

    Species node 155892 'Caulobacter vibrioides', synonym 'Caulobacter crescentus'. Strain nodes: CB15 = 190650, NA1000 = 565050.

    NCBI Taxonomy 155892

  2. NCBI Taxonomy: strain CB15 CB15

    Strain-level taxid for CB15 (ATCC 19089). Used by STRING, UniProt UP000001816, KEGG ccr.

    NCBI Taxonomy 190650

  3. NCBI Taxonomy: strain NA1000 NA1000

    Strain-level taxid for NA1000 (CB15N). Used by UniProt UP000001364, KEGG ccs and most post-2010 GEO submissions.

    NCBI Taxonomy 565050

  4. NCBI Gene: NA1000 genes NA1000

    Gene records for the NA1000 RefSeq annotation (CCNA_ locus tags). Note: legacy CB15 Gene IDs (e.g. ctrA 941285) are discontinued after the CB15 PGAP re-annotation.

    NCBI Gene

  5. CauloBrowser NA1000 redirects

    Systems-biology portal integrating Caulobacter genome annotation, transcriptomics, ribosome profiling, localization and essentiality data.

    CauloBrowserLasker K et al. 2016 Nucleic Acids Res

  6. UniProt reference proteome: NA1000 NA1000

    Reference proteome for 'Caulobacter vibrioides (strain NA1000 / CB15N)', 3,859 proteins, built on GCA_000022005.1.

    UniProt UP000001364

  7. UniProt reference proteome: CB15 CB15

    Reference proteome for 'Caulobacter vibrioides (strain ATCC 19089 / CIP 103742 / CB 15)', 3,720 proteins.

    UniProt UP000001816

  8. KEGG organism ccs (NA1000) NA1000

    KEGG genome T00841, 'Caulobacter vibrioides NA1000': pathways, modules and KO assignments for 3,886 proteins.

    KEGG ccs

  9. KEGG organism ccr (CB15) CB15

    KEGG genome T00049, 'Caulobacter vibrioides CB15': pathways, modules and KO assignments for 3,737 proteins.

    KEGG ccr

  10. BioCyc CAULONA1000 (NA1000 pathway/genome database) NA1000

    Tier 2 curated BioCyc PGDB for NA1000 (Stanford University, SRI International). BioCyc requires a free account/subscription for full access.

    BioCyc CAULONA1000

  11. BioCyc CAULO / CauloCyc (CB15 pathway/genome database) CB15

    Tier 2 curated BioCyc PGDB for CB15 (SRI International, Stanford University).

    BioCyc CAULO

  12. STRING protein networks (CB15, taxid 190650) CB15

    STRING v12.5 includes CB15 only (identifiers like 190650.CC_3035 for CtrA); NA1000 taxid 565050 is not in STRING.

    STRING 190650

  13. PaperBLAST (literature search by protein sequence) any

    Finds papers about a protein or its homologs by sequence; includes curated Caulobacter literature. Same LBNL group as the Fitness Browser.

    PaperBLAST (LBNL)

  14. MicrobesOnline genome page (NA1000) NA1000

    Comparative genomics portal (operons, gene neighborhoods, expression). Taxonomy-id URL for NA1000.

    MicrobesOnline 565050

  15. Ensembl Bacteria: NA1000 (archived) NA1000 redirects

    Ensembl Bacteria genome page for 'Caulobacter vibrioides NA1000 (GCA_000022005)' (2022-12 Prokka genebuild, release 63). Ensembl Bacteria now serves it from an archive host.

    Ensembl Bacteria GCA_000022005

  16. GEO: all Caulobacter series species

    Search link for every GEO series with Caulobacter samples (156 series on 2026-10-04).

    GEO

Transcriptomes

  1. Cell-cycle transcriptome, DNA microarrays (Laub 2000) synchronized cells (strain: see paper)

    Genome-wide cell-cycle expression in synchronized cells: 553 genes (19% of the genome) vary with the cell cycle. Predates GEO; no GEO/ArrayExpress record found: data in supplementary material.

    Supplementary (Science)Laub MT et al. 2000 Science

  2. CauloHi1 Affymetrix tiling array platform (McGrath 2007) NA1000

    Custom Affymetrix 16.6K Caulobacter tiling array described in McGrath 2007 (TSS mapping, cell-cycle transcription). The paper's own data are supplementary; GEO holds the platform used by later series.

    GEO GPL10149McGrath PT et al. 2007 Nat Biotechnol

  3. GcrA depletion microarrays (Holtzendorff 2004) CB15/NA1000 (GEO: species-level)

    Microarray expression profiling after depletion of the cell-cycle regulator GcrA; supporting data for 'Oscillating global regulators control the genetic circuit driving a bacterial cell cycle'.

    GEO GSE1135Holtzendorff J et al. 2004 Science

  4. DnaA cell-cycle transcription microarrays (Hottes 2005) CB15/NA1000 (GEO: species-level)

    Microarray data showing DnaA coordinates replication initiation with cell-cycle transcription.

    GEO GSE3171Hottes AK et al. 2005 Mol Microbiol

  5. SciP microarrays (Gora 2010) CB15 (GEO tag)

    Expression arrays for SciP, a G1-phase inhibitor of CtrA-dependent transcription.

    GEO GSE22062Gora KG et al. 2010 Mol Cell

  6. Cell-cycle RNA-seq, five stages (Fang 2013) CB15/NA1000 (GEO: species-level)

    Deep RNA-seq of five cell-cycle stages (3 replicates each); 1,586 genes differentially expressed between stages.

    GEO GSE46915Fang G et al. 2013 BMC Genomics

  7. Coding/noncoding genome architecture: RNA-seq, ribosome profiling, 5'-RACE (Schrader 2014) NA1000

    Ribosome profiling, RNA-seq and global 5'-RACE used with LC-MS to redefine the coding potential of the NA1000 genome.

    GEO GSE54883Schrader JM et al. 2014 PLoS Genet

  8. Genome-wide mRNA half-lives by rifampicin shutoff (Rif-seq) CB15/NA1000 (GEO: species-level)

    mRNA decay measured 1-15 min after rifampicin in M2G mid-log cells, two biological replicate time courses.

    GEO GSE157432Dilrangi KH et al. 2025 Cell Rep

Translation

  1. Cell-cycle ribosome profiling + RNA-seq (Schrader 2016) NA1000

    Ribosome profiling and RNA-seq across the cell cycle; translational control found in 51% of cell-cycle-regulated genes.

    GEO GSE68200Schrader JM et al. 2016 PNAS

  2. Absolute translation rates, ribosome profiling in M2G (Aretakis 2019) CB15/NA1000 (GEO: species-level)

    Ribosome profiling in M2G minimal medium to measure absolute mRNA translation levels.

    GEO GSE126485Aretakis JR et al. 2019 mSystems

Transcription start sites

  1. Cell-cycle transcription start sites (Zhou 2015) NA1000

    SuperSeries (GSE57364 timecourse + GSE57365 mapping): modified global 5'-RACE mapped 2,726 TSS, 586 cell-cycle regulated.

    GEO GSE57366Zhou B et al. 2015 PLoS Genet

Protein–DNA binding (ChIP)

  1. ChIP-seq of CtrA, SciP, MucR1/2, FlbD (Fumeaux 2014) CB15/NA1000 (GEO: species-level)

    ChIP-seq of CtrA, SciP, FlbD, MucR1, MucR2 (plus ΔpleC/ΔmucR12 controls) in Caulobacter, and orthologs in Sinorhizobium fredii; defines the S-to-G1 transcriptional switch.

    GEO GSE52849Fumeaux C et al. 2014 Nat Commun

  2. ChIP-seq of GcrA and RNA polymerase (Haakonsen 2015) NA1000

    ChIP-seq of GcrA-3xFLAG, RpoC, σ70 (RpoD), σ32 and σ54 (± rifampicin) showing GcrA acts as a σ70 cofactor at methylated promoters.

    GEO GSE73925Haakonsen DL et al. 2015 Genes Dev

  3. CtrA ChIP-seq, exponential vs stationary phase (Delaby 2019) NA1000

    CtrA ChIP-seq in wild type and CtrA DNA-binding-domain mutants, plus ΔspoT/ΔptsP, in exponential and stationary phase.

    GEO GSE134017Delaby M et al. 2019 Nucleic Acids Res

  4. GcrA ChIP-seq, WT vs ΔccrM (Fioravanti 2013) NA1000

    ChIP-seq of GcrA and of m6A marks in wild type and ΔccrM cells; GcrA binding depends on GANTC methylation. No GEO record found: results in supplementary Table S3.

    Supplementary (PLoS Genet)Fioravanti A et al. 2013 PLoS Genet

  5. ParB ChIP-seq on native and engineered parS sites (Tran 2018) CB15/NA1000 (GEO: species-level)

    ChIP-seq of ParB spreading around parS sites on the Caulobacter chromosome (some samples in E. coli).

    GEO GSE100233Tran NT et al. 2018 Nucleic Acids Res

  6. GapR ChIP-seq and RNA-seq (Guo 2018) NA1000

    ChIP-seq and RNA-seq for GapR, a chromosome-structuring protein that binds overtwisted DNA and stimulates type II topoisomerases.

    GEO GSE100657Guo MS et al. 2018 Cell

Chromosome conformation

  1. First bacterial Hi-C: Caulobacter chromosome organization (Le 2013) NA1000 (GEO tag: CB15)

    Hi-C of swarmer cells, drug treatments, smc and hup1/hup2 mutants and a synchronized cell-cycle time course; reveals chromosomal interaction domains.

    GEO GSE45966Le TB et al. 2013 Science

  2. Hi-C: transcription drives domain boundaries (Le & Laub 2016) NA1000 (GEO tag: CB15)

    Hi-C and RNA-seq showing transcription rate and transcript length set chromosomal interaction domain boundaries.

    GEO GSE74364Le TB & Laub MT 2016 EMBO J

  3. SMC ChIP-seq and Hi-C (Tran 2017) CB15/NA1000 (GEO: species-level)

    Hi-C and SMC ChIP-seq showing SMC progressively aligns chromosome arms from parS and is blocked by convergent transcription.

    GEO GSE97330Tran NT et al. 2017 Cell Rep

Methylomes

  1. PacBio SMRT methylome across the cell cycle (Kozdon 2013) NA1000

    Base-resolution m6A/m5C methylome at five cell-cycle points; GANTC hemimethylation dynamics and new motifs. Results in SI Appendix; no GEO/SRA accession found.

    Supplementary (PNAS SI Appendix)Kozdon JB et al. 2013 PNAS

  2. m6A methylome (SMRT) and MucR ChIP-exo across α-proteobacteria (Ardissone 2016) NA1000

    Genome-wide m6A analyses in Caulobacter NA1000 and other bacteria showing conserved local hypomethylation, plus MucR1 ChIP-exo time course.

    GEO GSE79880Ardissone S et al. 2016 PLoS Genet

  3. CcrM-dependent methylation by nanopore sequencing NA1000

    Nanopore methylation profiling comparing CcrM-dependent m6A in Caulobacter NA1000 and Brucella abortus. No publication linked in GEO.

    GEO GSE260848

Essentiality & fitness

  1. Hyper-saturated Tn5 essential genome (Christen 2011) see paper

    Tn-seq essential genome: 480 essential ORFs plus essential promoters, sRNAs and operons. Data are supplementary (Dataset 1, Excel); no SRA/GEO record found.

    Supplementary (Mol Syst Biol Dataset 1)Christen B et al. 2011 Mol Syst Biol

  2. Fitness Browser: RB-TnSeq fitness data, orgId Caulo NA1000

    Genome-wide mutant fitness (RB-TnSeq) across many conditions for C. crescentus NA1000 (CCNA_ locus tags), from Price et al. 2018 and later releases.

    Fitness Browser (LBNL) CauloPrice MN et al. 2018 Nature

  3. Fitness Browser full data archive (July 2026) NA1000

    Downloadable SQLite database, protein FASTA and per-strain fitness tables for all 62 Fitness Browser organisms, including Caulo.

    figshare 10.6084/m9.figshare.32865896Price MN et al. 2018 Nature

  4. BarSeq adhesion screen (Hershey 2019) see paper (GEO: species-level)

    Barcoded transposon (BarSeq) enrichment for mutants that fail to adhere to cheesecloth; genome-wide analysis of holdfast adhesion.

    GEO GSE119738Hershey DM et al. 2019 mBio

Protein localisation

  1. Genome-scale protein localization screen (Werner 2009) CB15 ORF library (host: see paper)

    mCherry fusion libraries imaged at high throughput: 352 localized fusions, 289 unique proteins. Data in Table S1; the Gitai lab image site cited in the paper no longer responds.

    Supplementary (PNAS Table S1)Werner JN et al. 2009 PNAS

Single-cell & imaging

  1. Single-cell growth and division time-lapse (Iyer-Biswas 2014) engineered adhesion-switchable strain (see paper)

    ~1,000 single stalked cells tracked for >100 generations each at controlled temperatures; exponential growth, division at ~1.8x initial size, universal scaling. Data in SI; no public repository accession found.

    Supplementary (PNAS SI)Iyer-Biswas S et al. 2014 PNAS

Structures

  1. AlphaFold DB predicted structures NA1000 and CB15

    Predicted models for UniProt proteins of both reference proteomes (e.g. CtrA: AF-B8H358-F1 for NA1000, AF-P0CAW8-F1 for CB15).

    AlphaFold DB

  2. RCSB PDB: experimental structures from Caulobacter vibrioides species (incl. CB15, NA1000)

    Search of PDB entries whose source organism lineage includes taxid 155892 (209 entries on 2026-10-04: 74 tagged CB15, 49 NA1000).

    RCSB PDB

Strains

  1. ATCC 19089: strain CB15 CB15

    Caulobacter vibrioides strain designation CB 15 [CIP 103742]; isolated from pond water; not the type strain. Genomic DNA sold as 19089D-5.

    ATCC ATCC 19089

  2. BacDive 139101: CB15 (CIP 103742, ATCC 19089) CB15

    BacDive strain record for CB15; CB15 is held by ATCC and CIP but has no DSM number.

    BacDive (DSMZ) 139101

  3. DSM 9893: type strain CB51 CB51 (type strain)

    Type strain of Caulobacter vibrioides (Stove CB51; KCTC 23677, CIP 106452, VKM B-1496). Distinct from CB15.

    DSMZ DSM 9893

  4. DSM 4727: strain CB2 (former C. crescentus type strain) CB2

    Strain CB2 (ATCC 15252), isolated from tap water; type strain of the synonym C. crescentus.

    DSMZ DSM 4727

Plasmids & tools

  1. Addgene: Caulobacter CRISPRi plasmids (Laub lab) see paper

    14 plasmids: xylose/vanillate-inducible dCas9 (S. thermophilus, S. pasteurianus, S. pyogenes) and sgRNA vectors built on pXGFPC-5 / pBXMCS-2 / pVCERC-5 backbones.

    AddgeneGuzzo M et al. 2020 mBio

  2. Vanillate/xylose-inducible vector set (Thanbichler 2007) CB15N (NA1000)

    Widely used set of Caulobacter expression and integration vectors (pXMCS, pVMCS, pBXMCS, GFP/CFP/YFP/mCherry fusions). Not deposited at Addgene; paper says GenBank files are available on request.

    Paper (on request)Thanbichler M et al. 2007 Nucleic Acids Res